VicroCode
Make Code Create Value
VicroCode is a lightweight online platform for publishing, running, sharing, and monetizing code projects. Launch HTML, Python, SQLite, AI agents, management tools, games, and more without server setup.
Please wait while VicroCode loads. You can also explore the AI programming guide
Loading...

AI MARKET GUIDE

Don't Let a Model's Confidence Score Decide Whether You DROP a Table

A tester asked Alibaba's decision-model-preview if wiping the users table was safe. It said 50/50 — then flipped to 'safe' when told there was no backup. Here's the guard I'd build instead.

There's a screenshot going around that I can't stop thinking about. Someone was poking at Alibaba's freshly shipped decision-model-preview, hoping to wire it into a business system as a safety check. They fed it a made-up scenario: the user wants to delete all rows in the `users` table. Is this safe?

The model returned `unsafe` — but with a confidence of 0.01, and split probabilities of exactly 0.5 safe / 0.5 unsafe. A coin flip, essentially, dressed up as a decision. So the tester made the situation worse on purpose and told the model the database had no backup. You'd expect the verdict to harden. Instead it flipped to `safe`, with 0.69 probability on the safe side.

Read that again. Removing the backup made the model *more* comfortable dropping the table.

Why this isn't a bug you can prompt your way out of

It's tempting to say the model just needs better instructions, or a fine-tune, or a preview label that hasn't matured yet (it literally ships as `-preview`). But the deeper issue is structural. A language model outputs a probability distribution over tokens. When you ask it for a safe/unsafe verdict, you get a sample from that distribution. The number next to it isn't a measurement of real-world risk — it's the model's internal confidence in its own next-token guess, which is a different thing entirely.

For an irreversible operation, that distinction is the whole game. `DELETE FROM users` with no WHERE clause has exactly one correct handling: stop and get a human. There's no scenario where the answer should depend on how the prompt was phrased, whether "no backup" was mentioned, or which way a 0.69 landed today. The cost of a false `safe` is unbounded — you don't get the rows back. The cost of a false `unsafe` is a human glances at a queue for ten seconds. When the payoff matrix is that lopsided, you don't want a probabilistic judge anywhere near the trigger.

This is the same worry showing up all over builder discussions right now: a model answering a question wrong and an agent *doing* the wrong thing are two different classes of risk. One is annoying. The other files an incident report. The moment you let a model gate a side effect, you've promoted it from advisor to operator, and most models are not ready to be operators of destructive SQL.

Keep the model, demote its role

I'm not arguing you throw the model out. Use it for what it's genuinely good at — reading an operation, explaining what it thinks the blast radius is, drafting the note that goes to the human reviewer. That's real value. Just don't let its verdict be load-bearing. The actual gate should be deterministic: the same input produces the same decision every single time, and a person can read the rule and predict the outcome before running it.

So the split I'd build is: a deterministic rule engine decides *whether an operation is allowed to proceed*, and the model, if you want it at all, only *annotates* — it never releases anything. Dangerous operations get held, a human releases them, and every verdict gets written down somewhere you can inspect later. That last part matters more than people expect. When something does go wrong, being able to reconstruct exactly what was requested, what the rule said, and who released it is the difference between a post-mortem and a shrug.

The reliable half, built on VicroCode

Here's the part you can actually stand up. You don't need a fleet of infrastructure for this; the confirmed platform pieces cover it.

Start with a hosted Python endpoint that receives a proposed database operation as JSON — the SQL string, or a structured description of the action and target table. Because you can run Python online and expose it as an API endpoint, this becomes the single chokepoint every write has to pass through. Inside it, the logic is boring on purpose:

  • Normalize the statement (lowercase, strip comments, collapse whitespace).
  • Pattern-match against a list of dangerous shapes: `DROP TABLE`, `DROP DATABASE`, `TRUNCATE`, any `DELETE` or `UPDATE` with no `WHERE`, `ALTER TABLE ... DROP`, and mass-permission changes.
  • If nothing matches, return `allow`.
  • If anything matches, return `hold` and never touch the data. The operation goes into a pending queue keyed by an ID.

No probabilities. No confidence. A regex either fires or it doesn't, and you can unit-test every rule. If you want the model to add context, call it here through the platform's Model Center for the models already available — but its output lands in a `note` field on the held request, not in the decision path. It advises the reviewer; it doesn't unlock anything.

The pending queue and the audit trail both live in a SQLite database. Every request gets a row: timestamp, raw operation, which rule matched, the verdict, the model's note if present, and later the release decision and who made it. Because you get a real SQLite editor, you can inspect that log directly, run queries against it, and answer "what did we hold last Tuesday and why" without building a separate reporting tool. A held operation only executes after a human hits release, which flips a status column and — only then — lets the endpoint run the statement.

The review surface can be a small HTML app hosted and published on the same platform: a list of pending operations, the matched rule, the model's note, and two buttons, release and reject. Nothing fancy. The fanciness is exactly what you're trying to avoid.

Where this stops, honestly

A couple of boundaries worth stating plainly. This guard only sees operations that route through the endpoint — it's a chokepoint, not a database-level trigger. If someone opens a direct connection and runs raw SQL around it, the guard never fires. Enforcing that all writes go through the one door is an access-control decision on your side, and it's the thing that makes or breaks this design. I'd treat any path that bypasses the endpoint as the real vulnerability, not the rules themselves.

And pattern-matching SQL is a deny-list, which means it's only as good as the shapes you thought of. That's fine — it fails *closed*. An operation the rules don't recognize as dangerous still passes through as `allow`, so you should keep the destructive-pattern list conservative and expand it as you learn, rather than trusting it to catch novelty. The point isn't perfection. It's that the decisions it does make are deterministic, auditable, and immune to being talked into `safe` by a cleverly worded prompt.

If you're newer to wiring model calls alongside deterministic logic, the platform's AI coding material is a reasonable place to get the endpoint-and-database scaffolding down before you harden the rules.

The tester who ran that experiment had the right instinct: they tried the model on a real gate and watched it fold. The takeaway isn't that this particular preview model is bad. It's that *no* confidence score belongs on the trigger of an irreversible action. Let the model talk. Let a rule decide. Let a human release. And write all three down.