Research
Beyond the lethal trifecta
The lethal trifecta has become the default model for reasoning about AI agent risk, but it only captures the confidentiality corner. As agents gain tools, integrity and availability risks demand a fuller model built on data flow.
Yusuf Hussein ·
Field notes
The Meta AI security incident shows the risk of unauthorized agentic workflows
Attackers reportedly used Meta's AI support assistant to take over high-value Instagram accounts by combining legitimate capabilities into an unauthorized workflow. The broader lesson is that tool-calling AI agents introduce security risks at the workflow level.
Thomas Vissers ·
Case study
How we helped a leading European bank secure their financial AI assistant
Blue41 tested how a European banking AI assistant handled untrusted transaction data. The case shows why financial institutions need runtime visibility and layered controls when deploying AI assistants in banking.
Thomas Vissers & Tim Van hamme ·
Field notes
Lessons from the Blue41 x OWASP workshop on breaking AI guardrails
Tim Van hamme ·
Research
Inside the AI security conclave
At last week's invite-only Dagstuhl Seminar on LLM Security, Blue41 joined leading researchers from academia and industry in reaching a clear consensus: AI security must evolve beyond input-output filtering to systems-level defense.
Tim Van hamme ·
Research
Why AI security must move beyond guardrails
Researchers from OpenAI, Anthropic, Google DeepMind, and ETH Zürich published a joint study with a sobering conclusion: the guardrails protecting today’s AI applications can be bypassed across the board.
Thomas Vissers ·
Research
Jailbreaking: The Researcher's Playground
Jailbreaking: Not just a researcher's game. Discover why today's AI safety measures might be more fragile than you think.
Tim Van hamme ·
Explainer
Unmasking Prompt Injection Attacks
This article explains what (indirect) prompt injection attacks are, how they differ from jailbreaks, and offer key strategies to secure LLM-powered systems.
Tim Van hamme ·
Case study
Real-world attacks on LLM applications
Attacks like prompt injections, indirect prompt injections, and the inherent issue of hallucinations are not just theoretical threats—they are being actively exploited and cause damage to businesses.
Thomas Vissers ·
Explainer
Hallucinations: LLMs' major reliability problem
Despite the powerful capabilities of LLMs, enterprises face challenges when moving from prototypes to production-grade AI deployments. A key concern when exposing LLM-based systems to the real word is reliability, most notably exemplified by so-called hallucinations.
Thomas Vissers ·
Field notes
LLMs: The Next Frontier in Cybersecurity?
This post delves into the intricacies of LLMs, exploring how they work and the vulnerabilities they present, especially when used in applications such as AI assistants.
Jo De Brabandere ·