AI Reviews That Read Nothing — Thirty-Four Verdicts in a Hundred
By Vladimir Mikhalev · Senior DevOps Engineer · Docker Captain · CNCF Ambassador · IBM Champion
Thirty-four of a hundred. That is how many of the images my fleet pins would have had their next AI code review written against nothing. The AI reviewer that judges every dependency update by its release notes was resolving them to a repository that did not exist. It returned a verdict anyway. I found it on 20 September by resolving each pinned image the way the reviewer resolves it, then asking GitHub whether the repository it named exists.
The model was not the part that failed. Every review that met a missing repository said so, in its last section: the release notes could not be retrieved. The line at the top said NEEDS ATTENTION, and the top line is the one a person reads. The shape is older than any model. A reconciliation can balance perfectly against the wrong account, and a change record can describe a different change: every figure correct, every figure about something else. A language model does not change that arithmetic. It makes the wrong answer longer and better written. One layer down, in the checks themselves, the same failure is the health check that had never been able to fail.
What an AI review of nothing looks like
Before any version bump moves, a model reads the upstream release notes against the compose file that pins the image and answers on one line: SAFE TO APPLY, NEEDS ATTENTION, or DO NOT APPLY UNATTENDED. To read the notes it needs a repository, so something has to turn an image name into one. That something was a short table of exceptions, and past the table, convention.
# review-source.sh: the convention past the table, image name in, repository outimage="${1#docker.io/}"case "$image" in ghcr.io/*|quay.io/*|lscr.io/*) echo "none" ;; # these registries name no GitHub owner */*) echo "$image" ;; # henrygd/beszel-agent becomes henrygd/beszel-agent *) echo "none" ;; # an official image has no owner to guessesacConvention is right often enough to look right, and it fails in the direction that produces an answer rather than an error. Renovate and Dependabot do not guess: they read the address the publisher wrote into the image, the standard source label in its configuration, and keep a list of exceptions behind it. That is the right order, and it is not an answer either. I read that label from the registry for the images in this fleet’s table. Thirteen named the right repository. Thirty carried no label at all. Five pointed at the repository that builds the image rather than the one that writes the software, which is the same wrong project as a redirect, written down by the publisher. A label is a claim. Something still has to check it.
The beszel agent is published under an image name that looks like a repository of its own, and no repository has that name. The agent ships from the hub’s own repository. Dify publishes four images from one repository and moves them together, so on 10 September a single Dify release produced four reviews, four not-found answers, and four verdicts of NEEDS ATTENTION that carried no information about the upgrade at all.
The fix is to stop trusting any name, guessed or declared, and prove it with the one fact the fleet already holds: the version it pins.
#!/usr/bin/env bash# verify-mapping.sh: a mapping is accepted only when the repository exists# under exactly that name and carries the version the fleet pins, as a tag.# Exit 0 proven, 1 disproven, 2 GitHub did not answer, which is no verdict.set -euo pipefailrepo="$1"pin="$2"api="https://api.github.com/repos/$repo"auth=(-H "Authorization: Bearer ${GITHUB_TOKEN:?}")lower() { tr '[:upper:]' '[:lower:]' <<<"$1"; }body=$(mktemp)trap 'rm -f "$body"' EXIT
code=$(curl -sSL -o "$body" -w '%{http_code}' "${auth[@]}" "$api" || true) # a transport failure prints 000case "$code" in 200) ;; 404) echo "no repository at $repo"; exit 1 ;; *) echo "GitHub answered $code for $repo: no verdict"; exit 2 ;;esacname=$(jq -r '.full_name' "$body")if [ "$(lower "$name")" != "$(lower "$repo")" ]; then echo "$repo redirects to $name: map to $name only if that is the project" exit 1fi
for tag in "$pin" "v$pin"; do code=$(curl -sS -o /dev/null -w '%{http_code}' "${auth[@]}" "$api/git/ref/tags/$tag" || true) case "$code" in 200) echo "proven: $repo carries $tag"; exit 0 ;; 404) ;; *) echo "GitHub answered $code for $repo tag $tag: no verdict"; exit 2 ;; esacdoneecho "$repo exists but has no tag for $pin: mapping unproven"exit 1A repository that answers and carries your version is evidence. A repository that merely answers is a coincidence, and coincidences scale. A GitHub that does not say yes or no, on a rate limit, a dead token or no network at all, has said nothing, so the script exits 2 rather than report an absence it never saw. Ask GitHub for a repository called mysql under the mysql account and it answers, through a redirect, with one called mysql-docker: a real project, the packaging and not the database. It was one of four plausible names that failed the tag test and never went in. A mapping that points at the wrong project is worse than none. It produces confident notes about somebody else’s software.
I have been writing glue like this since it resolved hostnames instead of source trees, and this failure has never once announced itself where anyone was looking.
The table has a history, and the history is the finding. The first fix, on 10 September, was twenty-six mappings written by hand, and the commit that added them said every image whose notes live elsewhere now had an entry. Ten days later, with eleven more images in the fleet, the sweep found thirty-four of a hundred pointing at nothing, so at least twenty-three of them had been pointing at nothing on the day the list was called complete. The thirty-four came apart into three kinds. Ten were missing mappings, each accepted only after the pinned version turned up as a tag. Seven were publishers that do not use GitHub at all, moved to a table of their own with the reason beside each. The rest were official images and registries that name no owner, for which convention had never produced an address in the first place. The first fix was a list. The second came with a test, and with a fourth answer: not found is a finding about the lookup, and the record now says so in a field of its own, beside a top line that still has only three values to choose from. That top line is the next thing to change.
Why this keeps happening
Four mechanisms stack, and each one turns a missing thing into an ordinary-looking thing.
A missing address answers like an empty one. GitHub returns 404 for a repository that does not exist and, by its own documentation, the same 404 for a private one your credential cannot see, so that it never confirms a private repository exists. The reviewer’s fetch failed, the prompt said the notes were not available, and the model did exactly what it was told. It said so, and listed what to check by hand. A publisher that does not use GitHub at all is a third answer. Atlassian, Microsoft’s SQL Server, GitLab Runner and Forgejo publish their notes elsewhere, and the fleet now says so from a table instead of reporting a failure it was always going to have. A rate limit is a fourth, and it is the one that publishes a smaller world. On 26 September the fleet’s own catalog was refused a third of the way through and published what it had. For an hour that was a fleet with a third of its restore scripts.
Nobody opens the envelope. A weekly job asks a model which of the week’s commits carry a rule worth keeping, and opens an issue with the answer. One week the issue arrived with its table of commits and nothing said about any of them. The response had come back over HTTP 200 with a stop reason of max tokens, 44,175 tokens in, 4,000 out, and not one text block. On the model it uses, Claude Sonnet 5, thinking is on by default, and Anthropic’s documentation is explicit that the token ceiling covers thinking and answer combined. The whole budget went into reasoning, and the answer never started. Joining an empty list of text blocks yields an empty string, which prints, evaluates as false, and writes to a file without complaint. The stop details field, the natural place to look, is filled in only for a refusal; for a truncated answer it is null. The fix was not a bigger ceiling. The job turns thinking off, raises when no text comes back, and says in the report itself when an answer was cut short.
The reviewer reads a description of the software, not the software. On 22 September dashy’s release notes said the default user was changing from node to 1000, and the review called it NEEDS ATTENTION, correctly from the notes alone: a changed container user over a bind-mounted file is how an upgrade breaks quietly. The registry disagreed. Version 4.7.5 declares the user called node, 4.7.7 declares 1000 for user and group, and both resolve to the same user and group. A rename, with nothing to change anywhere. A release note is a sentence somebody wrote. The image is the thing that runs. Since that night the review receives both image configurations from the registry, and where they disagree with the notes about the user, ports, volumes, entrypoint or health check, the configurations win. Silence is a description too: a game server’s release arrived marked breaking with no notes at all, and its diff was two lines of Dockerfile plus a new variable that decides whether the spectator relay forwards player voice. The marker said breaking. The diff said privacy.
An automation reads a summary instead of the jobs. The job that refreshes pinned image digests pushed a change, read its CI, and reverted it. The change was good. The stack came up, the HTTPS smoke test passed, all three Trivy scans passed. The only red job was the freshness check, a designed alarm that goes red whenever any pin in the repository lags upstream and says nothing about whether the digest just written boots. The automation read the run’s overall conclusion, and a conclusion cannot tell a designed alarm from a failure; over fourteen days, 254 of 281 red runs across the fleet had no failing job except that alarm. A model in the triage seat would have made the revert eloquent. It would not have made it right.
Four failures, one shape: a lookup, an envelope, a description and a summary, each standing in for the thing itself. Published research has the same shape from the model’s side. Google’s work on retrieval found that frontier models, handed insufficient context, answer wrong rather than abstain. Atlassian measured its production reviewer and found 22 percent of comments not grounded in what the reviewer was given. In August a team at Ionix watched their reviewer lose read access to the repository and keep producing confident, well-formatted reviews of code it had never read a byte of. None of those counted what this fleet counted, a verdict whose source was never fetched, and every one of them points the same way.
A pipeline that cannot tell an empty answer from no answer will eventually publish the second one.
Risk and blast radius
Direct exposure is the verdict somebody acts on. A third of the reviews told a person to look hard at an upgrade nobody had read, which is how a human was called in for beszel on 20 September. Raise the alarm on a third of all bumps and everyone downstream learns that the alarm means nothing, which is the lesson a review exists to prevent.
Systemic exposure is the gate. The review runs in front of every version bump of every image the fleet pins, a hundred of them that week, and one resolver fed every one of those reviews. A defect in a gate is the same wrong decision applied to everything that passes through it, at machine speed, with a paper trail that looks like diligence. This fleet is one maintainer’s. Put the same resolver in front of a thousand engineers and the failure keeps its shape. What grows is the number of people who learn to ignore it.
The code in this fleet comes from the same kind of machine, and the rule holds for it too: nothing an AI agent says about its own work counts as evidence. A fleet conformance checker was patched, re-run, and reported every repository meeting the standard, and the next day a fresh clone showed the patch had never been pushed. The clean number came from code that existed nowhere. The evidence page keeps the agent’s mistakes beside what caught each one.
Assurance exposure is the one that reaches the board, and it is quiet. SOC 2’s change management criterion asks whether changes are documented, assessed and approved through a structured process with version control. ISO 27001’s change management control asks whether changes to information systems are planned, assessed, authorized, tested and documented. Both are satisfied by an assessment that happened. Neither asks what the assessment was holding, and no change record I have audited has had a required field for the inputs of the assessment, only for its outcome. An AI verdict pasted into a change record satisfies every word of both while describing an upgrade it never read. OWASP’s Top 10 for LLM applications files this under misinformation, its ninth entry, and names the cure as grounding and verification, which is a standard a CISO can cite and a vendor cannot wave away. Documents rot the same way. Every security policy in this fleet told a reporter to fetch an encryption key from a page that did not exist, in ninety-seven repositories, from April until 25 September, and none of the checks that read the fleet ever opened the link.
Since 24 September every release in this fleet carries a git archive, a keyless signature and SLSA provenance, a statement of what was built from what, which anyone can check without trusting the repository. The review of the version bump inside that same release is a paragraph in its notes. We learned to ask a binary where it came from. Ask the verdict what it read.
The shape of the answer already exists. Provenance has a subject, the thing attested; a predicate, what was done to it and from what; and a builder. A verdict has the same three parts: the image at a version, the documents the reviewer read with a hash of each, and the model, its version, its token counts and how it stopped. Write the record in that shape and a verdict becomes something a second system can check, not a paragraph a person has to believe. Nothing on the market publishes it yet. Dependency bots attach the release notes they found; no AI reviewer I have evaluated attaches what it read.
The same question travels to the review of a merge request without changing a word. Which files were in the context window, and was the diff cut to fit it? A reviewer that cannot say has reviewed the part of the change it happened to see.
So the question to put to any automated review, yours or a vendor’s, stops being how good the model is. The questions about confidentiality, what leaves your perimeter and for how long, you ask already. These five are the ones nobody asks, and a team that buys a reviewer instead of building one has nothing else to hold it with:
- Which document did it read for this verdict, and at which version?
- Show me the last ten verdicts whose source was missing, and their top line.
- In the record, not the logs: can it tell a refused answer, a cut-off answer and an empty one apart?
- What can the verdict change by itself, with nothing but its own word behind it?
- Whose credentials does it spend, for how long are they valid, and could anything outside its pipeline spend them?
A system that can answer all five in one click has a reviewer. A system that cannot has a writer.
The failure modes, by what the reviewer was holding
Sorted by what the reviewer was holding when it ruled, the failures above come to nine: nothing, the wrong project, a blank, a description, silence, a third of the fleet, a summary, another workflow’s answer, and a local copy of code nobody had pushed. Each one is a dated finding on the ledger, with the rule it produced.
| What it was holding | What it looks like | What settles it | Seen here |
|---|---|---|---|
| Nothing | A verdict on an upgrade nobody read | Prove the source before the prompt, and name a missing one in a field | 34 of 100 images, 20 September |
| The wrong project | Confident notes about somebody else’s software | The repository must carry the pinned version as a tag, under its own name | mysql/mysql redirects to the packaging |
| A blank | A clean report with nothing in it | Raise on zero text blocks, with the token counts in the error | 4,000 tokens out, no text |
| A description | NEEDS ATTENTION on a rename | Give the model the registry’s image configurations, and let them win | dashy 4.7.5 to 4.7.7 |
| Silence | A breaking marker whose notes explain nothing | Read the diff, never the adjective | a game server’s privacy variable |
| A third of the fleet | A smaller fleet, published as the whole | Read everything before writing anything, and treat no answer as no verdict | 24 of 73 restore scripts for an hour, 26 September |
| A summary | A good change reverted | Read which jobs failed, and count an unreadable list as a failure | 254 of 281 red runs were one alarm |
| Another workflow’s answer | Twenty-three good changes reverted in one run | Judge a change only by the workflow that verifies it; skipped is not red | 23 reverts, 25 September |
| A local copy | A clean fleet report from code nobody pushed | Quote a number only after re-running from a fresh clone | the conformance checker |
Framework: prove the source, open the envelope, bound the verdict
Layer 1: prove the source before the prompt (week 1)
Every mapping from an image to a source repository is asserted, not assumed, and the assertion uses a fact you already hold. Take the address from the image’s own label where there is one, from a table where there is not, and run the mapping check above on every address, on a schedule: the label is the publisher’s claim and the tag is the proof. Keep the publishers that do not use GitHub in a short table with the reason written beside each entry. “Atlassian publishes release notes on its own site” is a different answer from “not found”, and it keeps the verdict about the image, not about a lookup that was never going to succeed. Then test the table the only way a table of names can be tested. Ask for every name, and prove the lookup can say no.
# tests/test-review-sources.sh: a lookup that cannot say no proves nothingrc=0./verify-mapping.sh henrygd/beszel-agent 0.20.0 || rc=$?if [ "$rc" -ne 1 ]; then echo "FAIL: a missing repository came back as exit $rc, not as absent" exit 1fi./verify-mapping.sh henrygd/beszel 0.20.0The fleet’s own versions are public. The reviewer carries its table of sources, with the reason written beside every publisher that keeps its notes elsewhere. The check above runs there as printed. And the test proves every mapping with a tag and plants all five answers: a missing repository, a redirect, a missing tag, the right repository, and a credential GitHub refuses, which has to come back as no verdict rather than as absence. Running that test across the table for the first time found four more mappings that pointed at a repository with no version tags at all, the packaging repositories of postgres, wordpress, xwiki and a game server. Real projects, with nothing in them for a reviewer to read, so every review of those images had been reading a blank. They moved to the table of publishers whose notes live elsewhere, with the address written beside each.
Owner: platform engineer.
Layer 2: open the envelope, then hand the model facts (week 2)
Treat every response as a container that may be empty. Check how generation stopped, check that text exists, and put the token counts in the error so the next person does not have to reproduce the failure.
# review.py: the envelope decides whether there is an answer at allimport anthropic
client = anthropic.Anthropic()
def answer(prompt: str) -> str: msg = client.messages.create( model="claude-sonnet-5", max_tokens=6000, thinking={"type": "disabled"}, # three verdicts need no reasoning pass messages=[{"role": "user", "content": prompt}], ) if msg.stop_reason == "refusal": raise RuntimeError(f"refused: {msg.stop_details}") text = "".join(b.text for b in msg.content if b.type == "text").strip() if not text: raise RuntimeError( f"no answer: stop_reason={msg.stop_reason}, " f"in={msg.usage.input_tokens}, out={msg.usage.output_tokens}" ) if msg.stop_reason == "max_tokens": text += "\n\n(Cut off at the token ceiling. Nothing below this was assessed.)" return textThen give the model something it cannot talk its way around. A release note is prose. The image configuration is a record the registry keeps, and it reads without pulling a single layer.
# registry-user.sh: what the image declares, read from the registry, nothing pulledfor v in 4.7.5 4.7.7; do printf '%s: ' "$v" docker buildx imagetools inspect "lissy93/dashy:$v" \ --format '{{ (index .Image "linux/amd64").Config.User }}' echodoneEach review in this fleet now writes a record beside its paragraph: the source it read, whether notes were found and why not, what the two image configurations differ in, the model, and the tokens in and out. That record is the provenance. The paragraph is the opinion. Every job here that asks a model anything goes through one shared helper, so the envelope is checked in one place and cannot drift between callers.
Owner: the engineer who owns the pipeline.
Layer 3: bound what the verdict can do (week 3 onward)
A review that only writes a paragraph is cheap to get wrong, so keep it that way. A verdict here can slow a change down. It can never wave one through. Inside a major version a bump proceeds on the CI-gated path. A verdict of DO NOT APPLY UNATTENDED reroutes it to a prepared branch for a person. When the review cannot run at all, the bump proceeds on the CI gate alone, as written. On 24 September a commit from a broken clone deleted the module every review imports, and for eighteen hours every review failed, sixteen of them before anyone noticed. The report said so on every line, and nothing shipped on a review’s word, because nothing ever does. The model’s credentials come from workload identity federation, so the only thing that can spend them is that workflow, on that branch, in that repository. The decision cost the ability to run a review from a laptop, which is the point.
The same rule covers any automation that is allowed to act, a reviewer or not: make the capability small rather than the request. A chat button that restarts a container is an endpoint on the internet that can act on a host. chatops-privilege-wall splits it in two.
# chatops-privilege-wall: the half that listens cannot actservices: edge: image: python:3.13-alpine read_only: true cap_drop: [ALL] networks: [edge-network] # Traefik joins this network; nothing else does volumes: - wall-queue:/queue # its one verb: write a request file
worker: image: python:3.13-alpine read_only: true cap_drop: [ALL] networks: [worker-network] # no ports, no router, not on the edge's network volumes: - wall-queue:/queue - /var/run/docker.sock:/var/run/docker.sock:ro
networks: edge-network: internal: true worker-network: {}
volumes: wall-queue: {}The edge writes a request file into a shared volume and holds no socket and no credentials. The worker holds the socket and has no inbound path at all. Each button carries a one-time token, and using any button burns every token issued for that alert, so there is no long-lived secret to leak: the token is the authorization.
Owner: SRE, with security owning the socket mount and the federation rule.
Tradeoffs
The cost is two or three GitHub calls per pinned image per run, six to eight registry requests per reviewed image, a seven-line table somebody keeps honest, a dozen lines of envelope checking, and a second container in a stack that had one. None of it is free, and the table is the part that rots. Ours rotted in ten days, which is why it now ships with a test.
What you buy is a verdict that names its source. For a hundred images that is a few hundred calls a day, inside one personal token’s hourly allowance of five thousand. The limit this fleet did meet was the catalog’s, on a night every job ran twice, and the rule that came out of it is the one the script above already keeps: a read that did not finish is not an answer. Reading one wrong verdict carefully enough to catch it takes longer than all of those calls together.
A few hundred requests a day. Thirty-four verdicts about nothing in a hundred. Take the trade.
The closing argument
The model was the most reliable component in that pipeline. It read what it was handed, said plainly what it could not read, and answered the question it was asked. Everything that failed happened before the prompt and after the response: in the lookup that chose the document, in the envelope nobody opened, in the summary read instead of the jobs. That is ordinary software, and none of it was tested until it failed, because a model was standing next to it.
Ask any automated reviewer to show you what it read, with an address, a version and a hash on it. If it cannot answer, you do not have a reviewer. You have a very fluent guess.
Sources
- Anthropic, Thinking (on by default on Claude Sonnet 5, and how to turn it off), accessed September 2026
- Anthropic, Steering thinking (the token ceiling covers thinking and answer combined), accessed September 2026
- Anthropic, Messages API reference (stop reasons, and stop details filled in only for a refusal), accessed September 2026
- GitHub Docs, Troubleshooting the REST API (404 instead of 403 for private repositories), accessed September 2026
- GitHub Docs, Git references (used above to prove a tag exists), accessed September 2026
- SLSA, Provenance, version 1.0, accessed September 2026
- OCI Image Format Specification, Annotations (the standard source label), accessed October 2026
- Renovate, Docker datasource (the source URL comes from the image’s label), accessed October 2026
- OWASP, Top 10 for LLM Applications 2025, Misinformation, accessed October 2026
- Tantithamthavorn et al., HalluJudge: reference-free hallucination detection for context misalignment in code review automation, 27 January 2026
- Joren et al., Sufficient context: a new lens on retrieval augmented generation systems, ICLR 2025
- Ionix, We built an AI PR reviewer, 13 August 2026
- AICPA Trust Services Criteria via ISMS.online, SOC 2 CC8.1, change management, accessed September 2026
- ISO 27001, 2022 edition, via ISMS.online, Annex A 8.32, change management, accessed September 2026
Discussion
If you run a model inside a pipeline and can produce, on demand, the exact document it read for its last verdict, drop a comment below with how. If you cannot, that is the more interesting comment. Counterarguments are welcome and the comment thread is where I respond first. For longer back-and-forth with senior practitioners, join the discussion on Discord.