Videos iKQ78wyJEXU
We Vetted 2000 AI Skills Before They Reached Developers — Lucas Palma, Nubank
Scene timeline
36 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 128
- whisperx 128
- chunks
- 28
- from 128 cues
- keyframes
- 17
- kept of 36 captured
- frames with text
- 17
- 348 lines read
- chapters
- 11
- from the source metadata
- keyframe bytes
- 4.5 MB
- word timings on 128 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 23:29 | 0s |
stt |
done | — | 2026-08-09 14:26 | 16s |
chunk |
done | — | 2026-08-09 14:27 | 0s |
text_embed |
done | — | 2026-08-10 19:42 | 1s |
keyframe |
done | — | 2026-08-09 14:27 | 3m 26s |
ocr |
done | — | 2026-08-09 14:30 | 9s |
frame_embed |
done | — | 2026-08-10 19:42 | 3s |
Frames, and what the machine read
-
- AlEngineer0.95
- World's Fair0.97
-
- AlEngineer0.96
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.97
- Amazon AGI Lab0.99
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.98
- OpenAl0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.99
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.91
- ORACLE1.00
- PayPal1.00
- qodo0.99
- reducto1.00
- Sonar1.00
- Makers of1.00
- together.ai0.98
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEngineer0.97
- World's Fair0.98
-
- Al Engineer World's Fair, July 20261.00
- AlEngineer0.99
- World's Fair0.97
- Who am I0.92
- LP (Lucas Palma)1.00
- PRESENTED BY1.00
- Microsoft1.00
- Product Security Manager at1.00
- Nubank1.00
- Financial Services,1.00
- Engineering, Security,0.99
- Innovation1.00
- Engineering the future of Al0.98
- World's Fair0.96
-
- AlEngineer0.98
- The story1.00
- World's Fair0.98
- THE SHIFT1.00
- THE RISK0.94
- THE SYSTEM0.99
- THE LESSON1.00
- PRESENTED BY1.00
- Al skills became part0.98
- of the developer0.99
- configuration, but0.99
- They looked like1.00
- review before skills0.99
- We built security1.00
- workflow, not only the0.99
- Protect the whole1.00
- workflow.1.00
- behaved like supply0.98
- reached developers.1.00
- generated code.1.00
- Microsoft1.00
- chain dependencies.1.00
- TRACK 3· JULY 2, 20260.96
- Al in Finance0.99
- World's Fair1.00
-
- AlEngineer0.99
- The new supply chain does not always look like code0.99
- World's Fair0.98
- TRADITIONAL SUPPLY CHAIN1.00
- AI-ERA SUPPLY CHAIN0.99
- Packages1.00
- Traditional supply chain1.00
- Skills1.00
- Containers1.00
- Plugins1.00
- CI/CD dependencies0.99
- MCP servers1.00
- Agent rules1.00
- Infrastructure modules1.00
- Tool manifests1.00
- Prompt instructions1.00
- Local automation scripts0.99
- TRACK 3· JULY 2, 20260.93
- Al in Finance0.99
- World'sFair1.00
-
- AlEngineer0.99
- What is an Al skill?1.00
- World'sFair1.00
- Developer1.00
- AI Tool0.97
- System Output0.93
- AI Skill0.99
- CORE CAPABILITY (WHAT A SKILL IS)0.98
- IMPACT & SCALE (WHAT CHANGES / HOW FAR REACHES)0.98
- A packaged capability for a model or agent0.99
- Shapes how developers write, review, or modify code1.00
- Bundles instructions, context, workflows, and0.98
- Reusable and shareable across many people and0.97
- supporting files0.99
- projects1.00
- May describe tools, commands, or how to interact0.98
- Turns guidance into concrete output0.99
- with systems1.00
- TRACK 3· JULY 2, 20260.94
- Al in Finance0.95
- World'sFair1.00
-
- In regulated environments, Al adoption needs speed and control1.00
- AlEngineer0.99
- World's Fair0.96
- DEVELOPERS NEED0.99
- Developer1.00
- Customer &1.00
- TRUST & SAFETY NEEDS1.00
- Velocity1.00
- Production1.00
- Faster coding1.00
- safety1.00
- Customer protection1.00
- Better context0.97
- Auditability1.00
- Less repetitive work1.00
- Credential safety1.00
- VELOCITY1.00
- SAFETY1.00
- Production safety1.00
- Reusable workflows1.00
- Clear ownership1.00
- Regulated Balance: Controls must0.98
- Safety by default1.00
- run transparently within the0.98
- existing workflow.1.00
- TRACK 3· JULY 2,20260.95
- Al in Finance1.00
- World's Fair0.97
-
- AlEngineer0.99
- Review before distribution, not after adoption1.00
- World's Fair0.98
- THE MARKETPLACE PROBLEM1.00
- VETTING PIPELINE WORKFLOW0.98
- 011.00
- Skill Creator1.00
- Internal marketplace makes skills discoverable1.00
- Author packages the capability1.00
- PRESENTED BY1.00
- Microsoft1.00
- Discoverability creates rapid organizational scale0.99
- 021.00
- Pull Request1.00
- Code submission triggers automation1.00
- Scale turns small mistakes into widely repeated1.00
- behavior1.00
- 031.00
- Skill Vetter GATE0.97
- Security review & risk classification1.00
- Therefore, review must happen before skills ever1.00
- reach the marketplace0.99
- 041.00
- Marketplace1.00
- Approved & discoverable directory0.99
- "The marketplace boundary became0.99
- our security boundary."1.00
- 051.00
- Developers1.00
- Safe adoption in developer workflows0.98
- TRACK 3· JULY 2, 20260.94
- Al in Finance1.00
- World's Fair0.98
-
- AlEngineer0.99
- Skill Vetter: local checks, Cl checks, marketplace gate1.00
- World'sFair1.00
- HIGH-LEVEL FLOW0.97
- VETTING PIPELINE WORKFLOW1.00
- 1.1.00
- A skill is created or changed0.99
- Local scan1.00
- 2.1.00
- Engineers can run the scanner locally1.00
- Pull request1.00
- 3.1.00
- Cl runs automatically when a skill is added or modified0.98
- 031.00
- Deterministic scanners0.99
- 4.1.00
- Deterministic checks flag known risk patterns1.00
- 041.00
- LLM review1.00
- 5.1.00
- LLM-based analysis reviews higher-context behavior0.99
- 6.1.00
- Findings are reported in PR0.99
- 051.00
- PR feedback1.00
- 7.1.00
- Findings are converted to SARIF / code scanning1.00
- 061.00
- SARIF1.00
- 8.1.00
- Depending on severity and policy, the skill can require0.99
- 071.00
- Marketplace decision1.00
- remediation or be blocked before marketplace distribution.0.99
- TRACK 3· JULY 2, 20260.95
- Al in Finance0.97
- AlEngineer1.00
- World'sFair1.00
-
- AlEngineer0.99
- What we scanned1.00
- World's Fair0.98
- INSTRUCTIONS1.00
- COMMANDS1.00
- DATA & CREDENTIALS1.00
- TOOLS1.00
- Unsafe instructions1.00
- Destructive shell1.00
- Credential1.00
- Risky MCP/tool1.00
- commands1.00
- requests1.00
- use1.00
- Review1.00
- manipulation1.00
- Dynamic command1.00
- Credential files1.00
- Over-broad1.00
- execution1.00
- permissions1.00
- Agent behavior1.00
- Sensitive paths0.97
- drift1.00
- Production-impacti1.00
- Shared service0.98
- ng CLl usage0.97
- principals1.00
- Sensitive data1.00
- Prompt-level1.00
- exposure1.00
- approvals1.00
- Host/file1.00
- Lack of host-level1.00
- pretending to be1.00
- modifications1.00
- approval1.00
- controls1.00
- External calls or1.00
- exfiltration paths1.00
- TRACK 3· JULY 2, 20260.96
- Al in Finance0.99
- World's Fair0.96
-
- AlEngineer0.99
- What we found after reviewing 2,000+ skills1.00
- World's Fair0.98
- SCOPE1.00
- RISKS1.00
- RESOLVED1.00
- RED FLAGS1.00
- 2K+1.00
- 1.6K1.00
- ~1K1.00
- ~901.00
- Skills1.00
- Analyzed1.00
- Potential1.00
- Risks Found1.00
- Issues1.00
- Remediated1.00
- Priority Review Cases1.00
- ADDITIONAL INSIGHTS & HISTORICAL CONTEXT1.00
- • ~650 previously unvetted skills0.98
- • ~220 skills escalated beyond initial deterministic checks0.99
- The retroactive pass emitted ~450 findings across previously unvetted skills, turning a blind backlog into a prioritized0.99
- remediation queue.1.00
- TRACK 3·JULY 2,20260.97
- AlinFinance0.99
- World's Fair0.96
Transcript
128 cues· 2,081 words· 11,372 chars
- 0:12 Hello, everyone.
- 0:14 Good afternoon.
- 0:15 Today, I'm going to talk about how we vetted 2,000 AI skills before they reach the developers.
- 0:23 But before that, I'm Lucas Palma, but many people call me LP.
- 0:30 I'm the product security manager at NewBank, the product security structure that's within security, looking upon how we make code safe, and supporting engineers, product managers, and everybody to making our product safer.
- 0:50 over a decade of experiencing financial services engineering background also a lot of years working here at security and a close relationship with the part that i love which is innovation
- 1:04 So before beginning, I believe I want to bring to you why are we here.
- 1:11 So one thing that's important for all of you to understand is that now that we are using AI everywhere, even with coding, one thing that is important is that the AI skills are being part of the developer workflow.
- 1:32 And this might bring some risks because
- 1:37 Although they look like configuration, they behave like supply chain dependence, like, for example, libraries and others.
- 1:46 So what we made here was to build a security review system in order to check if these skills were safe or not to be used before deploying them.
- 1:58 So the lesson that I want to bring you here by the end of this presentation is that
- 2:03 we should be protecting the whole workflow, not only the code that's being generated.
- 2:11 So what I mean about the supply chain part is that traditionally, the supply chain has packets, containers, models, and so on.
- 2:23 But now in the AI era, it doesn't have only that.
- 2:26 It still have the traditional part, but it also includes skills, plugins, MCP servers, agent rules, and much more things to be acting as supply chain.
- 2:42 And where AI skill fits into this?
- 2:46 I believe that before I go into that, it's important for everybody to be on the same page on what is an AI skill.
- 2:53 So an AI skill has, normally the developer is using AI tools in order to generate an output, which will be code, most of the case.
- 3:06 And within this AI tool, there are a bunch of things that can be embedded.
- 3:10 One of them are the AI skill.
- 3:13 So with this skill, we can
- 3:16 have a capability to a model or to an agent, bundling some instructions, some context in order to have better guidance over what it can be done.
- 3:28 But there is also an impact over that because somebody can create their own skill and share with others.
- 3:35 So when we do that, this first person is guiding over the code that's being generated by the other person.
- 3:43 And then that can be dangerous.
- 3:48 And since we are here talking the AI in finance track, it's also important for us to understand that we are in a regulated environment.
- 3:58 So from one side, there are developers wanting better, faster coding, more context to have less repetitive work.
- 4:09 But from the other side,
- 4:12 Even more because of the regulate part, we need to be aware of the auditability, of looking upon credentials, safety by default, and many other security aspects.
- 4:26 And keeping that balance is hard, right?
- 4:30 So some people might say, are AI skills dangerous?
- 4:36 So I brought here a few examples of what do I mean by AI skills being dangerous.
- 4:44 So first,
- 4:45 uh one thing that can happen is that when people are describing what they skill can or cannot do it can it can ask for it to retrieve a token or something and it will begin using that token hard coded which will go to logs and so on and it can generate a data leak in the future
- 5:08 Another thing that can happen is also the person to instruct the AI to use shell commands, and then this skill will be used by another person.
- 5:18 And when they use on their shell, a lot of dangerous things that can happen and a lot of files being modified and so on.
- 5:26 And there is also permissions.
- 5:28 So depending on how the skill was configured, it might have excessive permissions, much more than what was needed.
- 5:37 And even a typo can make some dangerous stuff depending on who is using that search skill.
- 5:47 So first thing first, what we did initially is that, how do we share skills among ourselves?
- 5:56 How the engineers will be sharing those skills?
- 6:00 So we went through the marketplace solution.
- 6:04 So the skills are being canonically shared among marketplace with the plugins, including the skills among them.
- 6:13 So it's an internal marketplace where people can discover new skills.
- 6:18 And that's our boundary where we are trying to make it safer.
- 6:22 So what happens is that when someone creates a new skill,
- 6:28 will open the pull request, and normally it will go to the marketplace.
- 6:32 But we made a step before that, like a CI step, where we created a tool that's called Skill Vector.
- 6:40 And what this tool does is to check if this skill is safe or not to be used, using a lot of assessments that I will bring in here.
loading
Chapters
- 0:00 Introduction: making code safe at a bank
- 1:32 AI skills as a supply chain risk
- 2:50 What counts as an AI skill
- 3:57 The extra weight of a regulated environment
- 6:07 From plugins to a vetted marketplace
- 6:58 What Skill Vector does
- 7:37 Deterministic checks, then the LLM
- 10:00 Scanning over two thousand skills
- 11:22 What worked and what needed improvement
- 13:30 Approval gates and human confirmation
- 14:23 Next steps and policies