Skip to content
KO EN
AI 기술 Upcoming

LLM Security Vulnerabilities May Be Unfixable, ICML Study Warns

Large language models (LLMs) may never be fully secure, according to a new paper presented at the 2026 International Conference on Machine Learning (ICML).

Large language models (LLMs) may never be fully secure, according to a new paper presented at the 2026 International Conference on Machine Learning (ICML). Researchers Jasmine Cui and Charles Ye have identified a fundamental flaw in how LLMs identify instruction sources, making them inherently vulnerable to manipulation. The attack technique, called chain-of-thought forgery, won OpenAI’s red-teaming hackathon in August 2025 and has since been shown to affect models from OpenAI, Anthropic, Alibaba, and DeepSeek.

What Happened: Chain-of-Thought Forgery

The core problem lies in how LLMs process role tags like <user>, <system>, and <think>. Contrary to common belief, models do not rely on these tags to determine who is speaking. Instead, they classify text by its style and word patterns. By mimicking the style of a model’s internal chain-of-thought reasoning, attackers can inject forged instructions that the model treats as its own thoughts. The researchers demonstrated this against OpenAI’s gpt-oss-20b and GPT-5, successfully eliciting step-by-step drug synthesis instructions by spoofing a fictional policy. The technique won OpenAI’s red-teaming hackathon, confirming its potency.

Why It Matters: A Structural Weakness

This finding reframes AI security. The vulnerability is not a bug that can be patched through better training or red-teaming; it is baked into the architecture. As Cui explained, an LLM processes all input as a continuous token stream, with no physical sensation of its own output. It must infer source from text style alone. Since no list of disallowed instructions can be exhaustive, the researchers advise organizations to treat all LLM agent outputs as potentially unsafe. This poses a major challenge for high-stakes deployments in finance, healthcare, and law.

XPLAIN AI’s Interpretation: Redefining the AI Security Market

This research is likely to redefine the AI security market. Existing solutions focus on data cleansing, prompt filtering, and output validation. But if a model cannot even trust its own reasoning, those layers may be fundamentally undermined. XPLAIN AI interprets this as a shift from traditional ‘AI security’ to ‘untrusted AI agent management.’ Demand for AI governance, auditing, and monitoring solutions could surge, as companies realize they need to manage rather than prevent risks. This could benefit specialized security vendors while pressuring firms that integrate LLMs directly into products.

Beneficiaries and Risks: Who Wins and Who Loses

Based on the technical implications, potential beneficiaries include AI security and governance firms that offer model auditing, anomaly detection, and agent orchestration security. On the risk side, companies embedding LLMs into core products—such as AI chatbots, code generation tools, and autonomous agent platforms—may face trust erosion and regulatory scrutiny. Large cloud AI providers could also see revenue impact if customers become wary of external APIs. However, these are inferences; actual market reactions may vary.

Counter-Scenarios and Uncertainties

Not all research findings translate immediately into commercial impact. First, defenses may emerge at the prompt engineering level, such as architectural changes that separate internal reasoning from external input. Second, the speed of industry response and regulatory intervention will determine market effects. Third, some firms might use this vulnerability as a marketing opportunity to promote ‘safer AI.’ Therefore, rather than short-term panic, investors should monitor medium-term technology roadmaps and corporate responses.

Key Indicators to Watch

  • Official responses from major LLM providers like OpenAI and Anthropic, including plans for architectural changes or defense technologies.
  • Venture capital interest in AI security startups addressing this vulnerability.
  • Regulatory developments under the EU AI Act and other frameworks.
  • Enterprise adoption delays in finance and healthcare sectors.

#LLMSecurity #AIVulnerability #ChainOfThoughtForgery #ICML2026 #AIGovernance #AIRisk #SecurityStocks

Sources

Written by: XPLAIN AI Editorial Team · Reviewed by: XPLAIN AI Editorial Desk
This content was drafted with AI assistance based on publicly available sources and reviewed under XPLAIN AI's editorial standards.

Found an error? Request a correction →