Yeah, exactly, the thinking process of the latest Claude models have really mucked up the responses they show to users, presumably to game benchmarks or something. My agents are constantly referring to things that happened behind the scenes, either in thinking or with subagents, as if I had full visibility into every aspect of everything they saw. But with Claude Code, everything is hidden.
I've taken to looking through the jsonl of sessions rather than trying to get Claude Code to explain what it means, and have better success about 50% of the time.
Older models work better, IMHO, and one can configure Claude Code to use any model that supports Anthropic Messages format, or a translator to other models, but the TUI itself is something I'm also straining against, and prefer Pi usually.
Anthropic's moat is actually workplace environments that are serving the opposing goals of 1) executive demands to use AI, and 2) legal demands to keep all company data on lockdown. In that environment, employees can get locked into whatever the approved AI methods are, and Anthropic excels in navigating that.
I've taken to looking through the jsonl of sessions rather than trying to get Claude Code to explain what it means, and have better success about 50% of the time.
Older models work better, IMHO, and one can configure Claude Code to use any model that supports Anthropic Messages format, or a translator to other models, but the TUI itself is something I'm also straining against, and prefer Pi usually.
Anthropic's moat is actually workplace environments that are serving the opposing goals of 1) executive demands to use AI, and 2) legal demands to keep all company data on lockdown. In that environment, employees can get locked into whatever the approved AI methods are, and Anthropic excels in navigating that.