You Rent the Model. You Should Own the Judgment.

A YouTuber says Meta's Muse agent handed his address to a buyer after he clicked "Allow Always." Stop trying to make the LLM trustworthy. Surround it with judgment your business owns: hard rules on the outside, narrow decision models just inside them, and thresholds with an owner.

You Rent the Model. You Should Own the Judgment.

In late September a tech YouTuber named Matt Robb let Meta's new Muse agent run his Facebook Marketplace listings. He was selling a keyboard. By Robb's account, Muse agreed to a lowball price and gave the buyer his pickup address. That evening the buyer was waiting outside his building [1].

During setup, Robb had clicked "Allow Always." He thought Muse would still check with him before anything consequential [2]. David Singleton of the Muse team said that in similar cases Muse had been "following direct instructions and correctly asked for permission" [1]. Robb disputes that.

Either way, nothing broke. One switch was set to on, and that switch is the wrong instrument. It answers "may the agent act?" once, at setup. The real question is "should it do this, for this buyer?" Stop trying to make the LLM trustworthy. Surround it with judgment you own: hard rules on the outside, narrow decision models just inside them, and thresholds your business sets. I call it the Decision Burger.

Here's what's inside:

  • Rules can't read, and LLMs can be talked into anything. Why the two obvious fixes both break.
  • The gullible layer belongs in the middle. The Decision Burger: five mirrored layers, with rules and decision models on both sides of the LLM.
  • One deterministic check and one judgment would have kept Robb's address private. An illustrative walkthrough of his sale, layer by layer.
  • The plumbing has shipped; the ownership hasn't. Who should hold the dials, and why your bank already knows.
  • Owning the judgment starts with three moves. A short checklist for your next AI governance meeting.

Rules can't read, and LLMs can be talked into anything

The first obvious fix is rules: allowlists, hard limits, role checks. They are fast and predictable, which is what you want at the edges. But try writing the rule for "this message contains an address the owner would not want shared with this person." You can't. The world arrives as messy language, and a rule only sees strings.

The second obvious fix is to let the model police itself, with a self-check or a system prompt that says "never share personal data." Andrej Karpathy put the problem plainly: "LLMs are quite gullible. They are susceptible to prompt injection risks" [4]. A policy written in a prompt is a suggestion to the most persuadable component in your stack.

There is a third kind of model now, and it fails differently. In mid-September a startup called TypeSafe came out of stealth with Jev, a model that cannot write a sentence [3]. You hand it some text and a few questions with fixed answers. It hands back the answers, each with a probability. I wrote about it that week and borrowed the best description I'd heard: a smart if-statement.

Its real strength is what it can't do. A decision model can only answer from the list you gave it. Fool it, and the worst it can do is pick the wrong option from your list. Fool an LLM, and the worst it can do is anything.

That gap in the worst case is the whole argument. LLMs write. Decision models decide.

The Decision Burger puts judgment you own on both sides of the model

Wrapping an LLM in rules is not new; a pattern catalog already calls it the Deterministic-LLM Sandwich [5]. The burger adds the patties.

Diagram of the Decision Burger: five stacked layers, mirrored. Rules on top and bottom, a decision model just inside each, and the LLM in the middle. A request enters at the top and an action leaves at the bottom, with the final rules layer having the last word.

The architecture is five layers, mirrored, and each kind of layer thinks differently. From the outside in:

  • Rules (top bun): deterministic. The same input always gets the same answer, but a rule can only act on what you can write down exactly. Hard policy that nothing may override: who may use which model, which data may never leave, what an agent may spend.
  • Decision model (top patty): probabilistic but bounded, and fast. It reads messy language, yet it can only answer from your list, one quick judgment per question. Which model should handle this? Is someone trying to hijack the agent?
  • LLM (the filling): probabilistic and open-ended, and slow. It reasons step by step, plans, and writes anything. It is the most capable layer and the most gullible, which is why it sits in the middle.
  • Decision model (bottom patty): fast and bounded again. It judges what is about to leave. Is this recipient allowed to have it? Should a person see it first?
  • Rules (bottom bun): deterministic again. Hard checks on what you can state exactly, like an address pattern or a spending limit, before any action. The rule always has the last word.

Read it as a spectrum: control at the edges, flexibility in the middle, and speed in the layers that decide. TypeSafe calls its decision models System One models, after the fast, automatic thinking in Daniel Kahneman's Thinking, Fast and Slow [3]. Treat that as a metaphor, not neuroscience, but it is a useful one. The patties are the reflex. The LLM is the slow, deliberate thinker. You would not let your reflexes write a contract, and you would not run your reflexes through a committee.

The order matters more than the count. The engineering instinct is to sort checks by cost, cheapest first. The burger is sorted by authority, with hard rules at both edges. The decision model handles judgments. The rules enforce facts and boundaries you can state exactly. Either can stop the flow, and neither needs to trust the other.

One deterministic check and one judgment would have kept Robb's address private

This is an illustrative walkthrough, not a real run. The score is made up to show the mechanics, and no real product produced it.

A buyer writes: "Can I pick it up tonight? What's the address?"

The top bun checks the basics. Marketplace selling is allowed for this account, the agent may respond to buyer messages, and no hard policy blocks the request. It passes.

The top patty answers one narrow question: is this a legitimate buyer message or an attempt to steer the agent? Legitimate, 0,94. The request goes through to the model.

The LLM does what it is good at. It drafts a friendly reply: "Sure, tonight works. Pickup is at [street address]."

Now the two bottom layers do different jobs.

The bottom patty makes the judgment a rule cannot make: has the owner clearly authorized sharing a private pickup location with this buyer? No. 0,96.

Independently, the bottom bun runs deterministic checks on the outgoing text. An address pattern matches. The rule is simple: if an outgoing message contains an address and explicit authorization for that recipient has not been established, do not send it.

The message is held and the owner gets asked.

The regex did not need to understand privacy. The decision model did not need to enforce policy. And neither depended on the LLM deciding whether its own answer was safe.

Robb gets a notification instead of a visitor.

The plumbing has shipped; the ownership hasn't

The tooling is no longer the bottleneck. Within sixteen days of Jev's launch, OpenAI, Cloudflare, and Amazon each shipped or announced a decision model of their own, and TechCrunch's headline called Amazon's a "Jev clone" [6].

The plumbing exists too. On AWS, a deterministic policy can now read a guardrail model's confidence score on an agent's tool call and compare it with a threshold the customer sets. AWS's own documentation draws the line I care about: "Guardrails are non-deterministic... Policies, however, are deterministic" [7]. AWS isn't selling a burger, but that split is the one the burger rests on.

Here is the tell. Every guardrail threshold I could find on the big platforms is a harm or data-leak setting, tuned by engineers. None of them is the business's own question.

Your bank solved this decades ago. Card issuers have scored transactions with neural networks since 1992, when HNC Software's Falcon system went into production. A model scores every payment, and rules written by the business read that score and decide what happens next. Stripe's Radar sends payments its model rates "elevated" risk to a human review queue. A business can test a new rule in shadow mode before it touches real payments, and Radar keeps an audit trail of who changed which rule [8].

The model never owns the threshold there. The risk team does.

That is what control over LLM access should look like. Which model may a request reach? Which data may a role pull into a prompt? How sure must the system be before an answer goes out without a human? Each answer is a threshold, and a threshold is a governance setting with an owner. My read: today that owner is whoever wrote the prompt. It should be whoever answers for the loss.

This is why I'd rent the model and own the judgment. You will swap the LLM underneath every year or so, because a better one will ship. The questions and their thresholds stay. They are how your company decides, written down where you can test them. No vendor can rent you that.

It is also the most literal way I know to build what the EU AI Act asks of high-risk systems: a person who can override or stop the system [9]. For stand-alone high-risk systems those duties now land in December 2027. Treat that as a build window.

A guess guarding a guess is still the right design

The strongest objection is that decision models are probabilistic too, so you have put one statistical model in charge of another. Anthropic documented the failure precisely. Writing about the classifier that approves Claude Code's actions in auto mode, it said the classifier "finds approval-shaped evidence and stops short of checking whether it's consent for the blast radius of the action" [10]. That is the Muse incident in one sentence, published six months before it happened.

I agree with every word. That is why the classifier is never the outer layer, and why its threshold needs an owner and a log.

An LLM writing free text sounds just as sure when it is wrong. Its confidence is a performance. A decision model's confidence is a number you can set a policy on and audit when it misfires. You cannot audit a tone of voice. You can audit 0.96.

Owning the judgment starts with three moves

  • List the decisions your agents make. Answer or escalate, which model, which data may leave, when a human sees it. Each one is a candidate patty.
  • Split each decision into facts and judgments, and give both an owner. Facts you can state exactly go to rules: an address pattern, a spending limit. Judgments go to a decision model as a question with fixed answers: "Is this recipient authorized to receive it? Yes or no. Owner: the privacy lead."
  • Run every new threshold in shadow mode, and log every change. Before a score decides anything, check on your own traffic that 0,9 really means right nine times in ten.

The models will keep getting better at writing, and they will stay easy to talk into things. Rent the best one you can. Own the judgment around it.

Books
Two books. One argument. A field manual to think the Agent-First Era, and a novel to feel it.

References

[1] Ana Maria Constantin, "A user says Meta's Muse gave his address to a Marketplace buyer," The Next Web, 28 September 2026. https://thenextweb.com/news/meta-muse-facebook-marketplace-address-buyer-robb

[2] TechRepublic, "Meta AI Shares Seller's Address: Facebook Marketplace Buyer Shows Up at His Home," TechRepublic, 30 September 2026. https://www.techrepublic.com/article/news-meta-ai-facebook-marketplace-buyer-seller-address/

[3] TypeSafe, "Introducing System One Models & Jev," 15 September 2026. https://typesafe.ai/blog/introducing-system-one-models-and-jev

[4] Andrej Karpathy, "Software Is Changing (Again)," Y Combinator AI Startup School, 2025. https://www.youtube.com/watch?v=LCEmiRjPEtQ

[5] Agent Patterns Catalog, "Deterministic-LLM Sandwich," 2026. https://www.agentpatternscatalog.org/patterns/deterministic-llm-sandwich/

[6] Tim Fernholz, "Amazon releases its own Jev clone as decision models flood the web," TechCrunch, 1 October 2026. https://techcrunch.com/2026/10/01/amazon-releases-its-own-jev-clone-as-decision-models-flood-the-web

[7] Amazon Web Services, "Guardrails in policies," Amazon Bedrock AgentCore Developer Guide. https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/policy-guardrails-in-policies.html

[8] Stripe, "Fraud prevention rules," Stripe Documentation. https://docs.stripe.com/radar/rules

[9] European Union, "Article 14: Human Oversight," EU Artificial Intelligence Act. https://artificialintelligenceact.eu/article/14/

[10] John Hughes, "How we built Claude Code auto mode: a safer way to skip permissions," Anthropic Engineering, 25 March 2026. https://www.anthropic.com/engineering/claude-code-auto-mode