Vulnerable AI Systems: Real Data, Responsible Design
29% of attacks bypass the security filters of the most widely used LLMs in production. It's not a bug. It's the nature of the system. LLMs are stochastic processes trained on human language—the most flexible, ambiguous, and manipulable medium that exists. That makes them incredibly powerful. And that's exactly why they're vulnerable. There's no patch for that. Only design. This talk presents the results of llm-break-bench: 3,360 adversarial tests on GPT-4o, Claude, Gemini, Grok, and DeepSeek using MLCommons AI Safety v0.5 and OWASP LLM Top 10 as standards. The numbers break intuitions. The smartest model in the benchmark is 5 times more vulnerable than the cheapest and 11 times more expensive. The most criticized by the press ends up second in security, and the reason behind that explains everything that's wrong with how the industry deploys AI today. The data is the starting point. The talk connects them to real use cases where LLMs are in production: RAGs, chatbots, agents, code assistants. It shows where design fails, what consequences it has (Air Canada paid for it), and how to build differently. The closing is actionable: 5 design pillars for AI systems that don't depend on the model for their own security, with real code from NVIDIA NeMo Guardrails and Meta LlamaFirewall. If you have an LLM in production or are about to, this talk changes how you design it.
Want to know more?
Join PyCon Colombia newsletter and get a complete overview of our events, speakers and community participation.


