The coming AI security crisis (and what to do about it) | Sander Schulhoff
Transcribed with NVIDIA Parakeet · Timestamps stay in sync with the live audio
I found some major problems with the AI security industry. AI guardrails do not work. I'm gonna say that one more time. Guardrails do not work. If someone is determined enough to trick GPT five, they're gonna deal with that guard. No problem when these guardrail providers say we catch everything, that's a complete lie. I asked Alex Kamarowski, who's also really big in this topic. The way he put it, the only reason there hasn't been a massive attack yet is how early the adoption is, not because it's secured. You can patch a bug, but you can't patch a brain. If you find some bug in your software and you go and patch it, you can be maybe 99.99% sure that bug is solved. Try to do that in your AI system. You can be 99.99%. I'm nine percent sure that the problem is still there. It makes me think about just the alignment problem. Gotta keep this god in a box. Not only do you have a god in the box, but that god is angry. And that god's malicious. That god wants to hurt you. Can we control that malicious AI and make it useful to us and make sure nothing bad happens. Today, my guest is Sander Schulhof.
This is a really important and serious conversation, and you'll soon see why. Sander is a leading researcher in the field of adversarial robustness, which is basically the art and science of getting AI systems to do things that they should not do. Like telling you how to build a bomb, changing things in your company database, or emailing bad guys all of your company's internal secrets. He runs what was the first and is now the biggest AI red teaming competition. He works with the leading AI labs on their own model defenses. He teaches the leading course on AI right teaming in AI security.
And through all of this has a really unique lens into the state of the art in AI. What Sander shares in this conversation is likely to cause quite a stir. That essentially all the AI systems that we use day-to-day are open to being tricked to do things that they shouldn't do through prompt injection attacks and jailbreaks. And that there really isn't a solution to this problem for a number of reasons that you'll hear. And this has nothing to do with AGI. This is a problem of today, and the only reason we haven't seen massive hacks or serious damage from AI tools so far is because they haven't been given enough power yet.
We already have this episode, so unlocking it is cheap and permanent: the full transcript, segment and word-level timestamps, and .txt, .srt, .vtt and word-level JSON downloads, as many times as you like, forever.
10 free credits is 10 episodes like this one a month. Episodes nobody has transcribed yet cost their length in hours, because we have to actually transcribe them.