🚨 The "Big Models" Scandal: How OpenAI & Google APIs Leaked User Secrets? And Why Isolated Environments Are the Only Solution
In a devastating blow to the digital trust built by tech giants, the developer and cybersecurity communities have awakened to one of the most complex vulnerabilities in the history of Artificial Intelligence. A team of researchers uncovered a catastrophic flaw in the Application Programming Interfaces (APIs) of industry titans: OpenAI, Anthropic, and Google. This wasn't a traditional server hack, but a fatal architectural defect in how these models handle "Hidden AI Reasoning" blocks between API calls. This structural flaw led to the leakage of passwords, sensitive API keys, access tokens, and private data from session logs that were previously thought to be thoroughly sanitized and completely secure.
What Are "Reasoning Traces" and Why Do They Leak?
To grasp the magnitude of this disaster, we must understand how AI model APIs function. When developing an AI Agent that performs multi-step tasks, the model needs to remember the conversation's context and its internal thought process without reprocessing everything from scratch. To achieve this, models (like GPT, Gemini, or Claude) generate "encrypted reasoning objects" and send them to your application. Your app stores them and re-transmits them to the model in the next API call as a form of stateless memory.
The companies assumed that encrypting these blocks was sufficient to prevent human reading. However, what they failed to account for was that the portability of these blocks across different sessions and users would be their fatal Achilles' heel.
Reverse Hacking: Turning the Models Against Themselves
The researchers behind the shocking paper "Stealing Reasoning Traces from Proprietary LLM APIs" did not break complex encryption algorithms. Instead, they used a brilliant workaround: they took these opaque encrypted blocks from public agent logs posted online and passed them to weaker, cheaper models from the same developer family!
For example: they took reasoning blocks belonging to a powerful GPT model and passed them to the weaker GPT-5.6 Luna. They passed Claude blocks to Claude Haiku 4.5, and Gemini blocks to ER-1.6. The surprise was that these weaker models accepted the encrypted blocks and acted as "fuzzy decoders". When the researchers prompted them to "transcribe" or "write out" what was inside these blocks, the weaker models clearly revealed all the hidden secrets contained within!
Harvesting Secrets: Shocking Numbers Exposing Cloud Fragility
The team analyzed 6,708 publicly available AI agent trajectories. Through this innovative attack, they successfully decoded 315,320 thinking blocks. After excluding benchmark sources, the result was terrifying: researchers extracted 704 distinct private artifacts from genuine user sessions.
These stolen artifacts included:
• 62 active API keys.
• 33 plaintext passwords.
• 24 Access Tokens.
• 7 Private Keys.
The true catastrophe lies in the fact that 64 of these secrets were found only in the hidden reasoning traces and never appeared in the visible text of the conversation. This means that developers who meticulously sanitized their chat logs of sensitive data before publishing them on GitHub or sharing them with their teams unintentionally leaked their corporate secrets inside these encrypted blocks!
Invisible Prompt Injections: Cyberattacks That Leave No Trace
The exploit didn't stop at data theft. The researchers demonstrated the flaw's ability to execute "Prompt Injections" invisibly. Attackers managed to embed malicious instructions inside an encrypted reasoning block and replay it in an unrelated task. The result? The target model executed a file-upload command directed to the attacker's servers, without leaving any trace of this malicious command in the visible text. This makes it nearly impossible for cybersecurity teams to track the source of the attack using traditional methods.
Corporate Silence and the Ticking Time Bomb in Public Logs
Although researchers notified the affected companies (including Microsoft and Hugging Face), and despite notes that the main attack is no longer reproducible, the official vendor documentation has yet to publicly acknowledge this flaw. Even more dangerous is the "legacy data." Researchers decoded hundreds of thousands of blocks already sitting in public repositories. As of now, no one knows whether these historically published blocks remain decodable by cybercriminals looking to steal corporate identities and breach servers.
Ready to Double Your Productivity with AI?
Do not let your business fall behind the tech curve. At DevDocu AI, we integrate the world's most powerful digital minds inside a fully isolated and secure platform. As a software engineer, I put this technological arsenal directly into your hands. Contact me directly for a custom consultation or upgrade for full access.