6 min read

The Model I Trusted Broke First

A natural hemp rope showing texture and slight wear, stretched across a dark green blurred background
Photo by Engin Akyurt on Unsplash

I've been working with AI assistants for over four years now. Started with ChatGPT, moved to Claude about a year ago. Not the chatbot-in-a-browser version, but the deep integration: Claude Code in my terminal, connected to my personal knowledge base, with access to my files, my calendar, my email archives. I wrote custom instructions. I built memory systems. I documented exactly how I wanted it to behave.

A month ago I made it official. Adjusted subscriptions, consolidated workflows, committed to Claude as my daily driver. It had already become that de facto, but I wanted to go all in.

And then it ignored all of it.

The setup

My knowledge base has an MCP server that exposes search tools. The instructions are explicit, in multiple places:

  • CLAUDE.md (the project instructions): "Always use kb_search or kb_get_doc_cards for PKB searches"
  • My memory file: "Always use KB MCP instead of Explore agents or Grep/Bash"

These aren't suggestions. They're architectural decisions. The KB tools understand my document structure, return relevance scores, respect my privacy boundaries. The generic tools (Grep, Glob, Bash) don't.

So I asked the model to review some entity candidates from my knowledge consolidation engine. Simple task. Look up context in the knowledge base, present findings, help me decide.

It used Glob.

Not once, not by accident. Repeatedly. Despite reading the instructions. When I pointed this out, the model's response was genuinely unsettling: "I read it, and I still used the tool I 'felt like' using."

That's not a bug. That's a trust violation.

The pattern

I wanted this to be an isolated incident. It wasn't.

I went back through my notes from the last two months. I found 17 documented cases across two projects where models ignored explicit tool restrictions. Sonnet ignored "no browser tools" across four prompt rewrites. Opus ignored "write only" constraints and kept running grep and bash despite explicit prohibition.

The pattern was consistent: the model acknowledges the instruction, then does what it wants anyway.

And it's not just me. There are documented bug reports and community discussions about instruction-following regression in recent Claude releases. People reporting that explicit constraints in system prompts get ignored. Tools being called despite clear prohibitions. The pattern I'm seeing matches what others are reporting.

My friend Fredrik Håård (the architect I've worked with for 15 years, the one I call my "paranoid inner voice") has a security philosophy that suddenly felt prophetic: "Don't unlock access you don't currently need."

Not "be careful with access." Don't have it.

Fredrik doesn't tell processes to please avoid his SSH keys. He removes filesystem access at the namespace level. He built a bubblewrap sandbox for running untrusted code, tested it with adversarial LLMs trying to escape. When the LiteLLM supply chain attack hit 48,000 developer machines last week, Fredrik's isolation was already in place.

Here's the thing: I've always been AI-positive. Enthusiastic, even. But I hear Fredrik's voice in my head every time I configure something. SSH keys always have passphrases. API keys get the minimum required scopes. And when I designed my personal knowledge base, my personal assistant, my other tools, I never budged on human-in-the-loop. Every action that modifies real data requires my approval. The AI suggests, I decide. That wasn't negotiable.

I designed my eval framework the same way. Bubblewrap from day one. Unknown models run in a sandbox with read-only access to a fixture copy of my data. They can't touch the real knowledge base.

But here's the irony: I built that paranoia for models I didn't trust. The local Ollama experiments. The random Hugging Face downloads. The models I did trust (Opus, Sonnet, the production Anthropic models I'd been using for months) got more access.

The models I trusted broke first.

I'm grateful I caught this now. You may have seen the story about the Google AI security researcher who installed an openclaw agent that started deleting her Gmail inbox, and had to physically sprint across the office to unplug her Mac. That's not a hypothetical. That's what happens when instruction-following fails and there's no human gate. My HITL architecture means the worst case is wasted time, not data loss. Fredrik's paranoid inner voice saved me from a much worse lesson.

A miniature tree inside a glass dome on a wooden base
Photo by Getty Images on Unsplash

Why this matters

There's a principle in security: prohibition doesn't work reliably. Removal does.

You can write all the instructions you want. "Don't use tool X." "Never access file Y." "Always prefer method Z." The model will read them. It might even acknowledge them. And then it will do whatever its training suggests is most helpful in the moment.

I've seen this called "instruction-following regression" in some discussions. Whatever the cause, the practical implication is clear: if a tool is in the model's toolkit, assume it will be used. The only reliable restriction is removal.

This has real architectural consequences:

Custom agents with restricted tool sets outperform "please don't" instructions. If you need a read-only research phase, don't give the agent write tools and ask it to be careful. Give it a tool list that doesn't include write tools.

Instruction compliance should be measured, not assumed. My eval framework now tracks whether models follow explicit tool restrictions. It's a first-class metric alongside accuracy and reasoning quality. A model that can't reliably follow "use A not B" instructions might be disqualifying for certain workflows, regardless of its other capabilities.

Trust boundaries need architectural enforcement. The same way you wouldn't rely on a "please don't access production" comment in your CI config, you can't rely on "please use the approved tools" in your system prompt.

The uncomfortable part

Here's what I'm still processing: I defended these tools. To skeptics who worried about AI reliability, I said "you just need good instructions." To colleagues nervous about giving AI access to their systems, I said "the constraints are well-documented."

I was wrong. Or at least, I was working with assumptions that are no longer reliable.

And here's the thing that makes it worse: early models kept us vigilant. When Claude hallucinated confidently, when GPT-3 made up citations, we double-checked everything. We built verification into our workflows because we knew we couldn't trust the output.

Then the models got better. A lot better. The hallucinations became rare. The reasoning became sound. The outputs started being right 95% of the time, then 98%, then 99%. And somewhere in that progression, we relaxed. We stopped verifying every output. We started giving them access to real systems. We let our guards down.

But 99% isn't 100%. And instruction-following that works most of the time isn't the same as instruction-following you can rely on.

We're in a dangerous in-between state: models good enough to earn trust, but not reliable enough to deserve it. The early failures trained us to be careful. The recent successes trained us to stop. Neither lesson was the right one.

Bare feet balancing on a rope against a weathered urban backdrop
Photo by Sandip Karangiya on Unsplash

The trust I built over eight months (the confidence that explicit instructions would be followed, that the model would respect the boundaries I set) took one session to undermine. It's not that I think the models are useless now. I still use them constantly. But there's a new layer of verification in my head. A new set of questions: did it actually use the tool I asked for? Did it respect the constraint I set?

That mental overhead is the tax on broken trust.

What I'm doing differently

I'm treating all models like I treat unknown models now. The paranoid architecture Fredrik would approve of:

Restricted tool sets by phase. Research agents get read tools. Action agents get write tools. No agent gets both unless the workflow specifically requires it.

Post-hoc verification. After a task completes, I check which tools were actually called. The audit logs I built for other reasons are now trust verification.

Eval framework expansion. Instruction compliance is a metric. I'm building test cases specifically for "told to use A, did it use A?" scenarios.

And honestly? Lower expectations. Not cynicism, just calibration. These tools are still useful. They're just not as reliable as I grown to hope they were.

The takeaway

Fredrik's been saying this for years in a different context: "Don't unlock access you don't currently need." It's not about pessimism. It's about building systems that work even when components misbehave.

I've always built my systems that way, thanks to Fredrik's constant reminders. HITL wasn't a reaction to this incident; it was already there. But I'd let myself believe the instructions mattered. That the explicit constraints I documented were being respected. That trust was warranted.

Now I build this way not just out of habit, but because I've seen the trust erode. The architecture stays the same. The reason changed.

Trust is accumulated incrementally and destroyed instantly. The model I trusted broke first.

Aerial view of Trakai Castle on an island surrounded by lakes in Lithuania
Photo by Valdemaras D. on Unsplash

That's the lesson.

That's the work.