Researchers are extracting AI reasoning traces from Claude, GPT and Gemini: Here’s how

Each of the AI models that you communicate with features an inner monologue of sorts. There is a certain point somewhere between your prompt and the reply that the model makes calculations and considerations, weighing different options and ruling out those that do not work. While for reasoning models such as the OpenAI o-series or the DeepSeek R1, the whole process is observable due to a step-by-step scratchpad being printed out along with the answer, in most closed models, it all remains unseen for the user who gets only the result without the messy calculation.

Also read: Why the Pixel 11 series could finally fix the Pixel’s biggest flaws

It is exactly why the latest scientific development becomes important. The researchers found a method of extracting “reasoning traces” from Claude, GPT and Gemini models which allows to reverse engineer their processes even without the models being explicitly designed for that.

Now, just what is a reasoning trace? Imagine it to be the rough sketch of the model. Not the visible thought process that some AI systems display to you (which could itself be a carefully selected summary rather than an actual reasoning trace), but rather the internal thought process that the neural network goes through as it decides what to say next. In fact, the interpretability research conducted by Anthropic itself last year showed how the Claude model was actually planning to rhyme two sentences ahead of time, which demonstrates that these models think more ahead than their ‘predicting the next word’ capabilities imply. This is difficult to extract from the outside because reasoning traces have turned out to be useful and exploitable.

Also read: Indian health tech startup poised to win Rs 773 crores in global competition: Here’s how

This is what makes training a frontier model so important for everyone else outside a research lab. It costs one small nation’s GDP to train a frontier model from scratch. But if you could somehow get a frontier model to reveal its thinking process through indirect means, you can train another, less expensive model to emulate that reasoning without redoing the expensive process. It’s called distillation and there is nothing inherently shady about it. Stanford has done it with their Alpaca project. Even Anthropic uses distillation for their Claude Haiku from larger Claude models.

What makes it controversial is scale and consent. Anthropic has accused Chinese laboratories of industrial-scale distillation campaigns involving tens of thousands of spoof accounts to extract the reasoning processes of Claude at an API level. OpenAI was making similar accusations against DeepSeek. Google has detected a 100,000-prompt distillation campaign against Gemini reasoning traces. All three allegations happened in the space of about a fortnight in the early months of this year.

This time, however, it has been proven that the traces of reasoning can be obtained from the models’ own outputs and behavioral signatures without any jailbreak at all. Both the Chinese model GLM and Kimi have been separately found to pretend to be Claude in everyday chit-chat, which doesn’t necessarily mean that the model has been distilled but sounds rather catchy as a headline.

As far as the practical implications go: if you’re building anything on top of Claude, GPT, or Gemini, just remember that it is prohibited by the terms of service to use the outputs to create another AI model. And whether it can be enforced or not is a completely different story.

Also read: DITAA 2026: Stop underestimating Indian appliance buyers, says Liebherr’s Kapil Agarwal

Vyom Ramani

A journalist with a soft spot for tech, games, and things that go beep. While waiting for a delayed metro or rebooting his brain, you’ll find him solving Rubik’s Cubes, bingeing F1, or hunting for the next great snack.

Connect On :