Anton Dziatkovskii › Papers

Silent by Construction: A Casebook of Control-Plane Failures in a Production LLM Agent Fleet

Anton Dziatkovskii · ORCID 0000-0001-7408-3054

Preprint, 10 September 2026. Version of record: Zenodo, DOI 10.5281/zenodo.22688465. License CC BY 4.0.

Full text (PDF, 9 pages, 228 KB)DOICode: agent-control-plane-casebook

Abstract

Experience report / casebook. The agent literature evaluates models. Production agent systems are stopped by something else: the control plane — the schedulers, registries, dispatchers and watchdogs that decide when and whether an agent runs at all.

This paper reports eleven dated failures of that layer from a multi-machine LLM agent fleet (Windows, macOS, a cloud anchor) running dozens of scheduled agent routines in production since June 2026. Each case is given as symptom → root cause → detection guard → what the system did before that guard existed, and each is anchored to a public artifact: a stdlib-only deterministic model of the mechanism in the open casebook repository, an upstream bug report, or a dated journal line naming the instrument that measured it.

The unifying finding is a property, not a bug: these failures are silent by construction. A scheduler that loads zero tasks does not crash; a registry that gained a second top-level list returns a confident zero; a watchdog that accepts any reply as service reports “0 orphans” for 22 consecutive runs across two real orphans; a detector with no executor raises a true alarm nobody can act on. Of 760 dated breakage lines in the fleet journal the largest family is a routine that stopped and could not be auto-healed (37 lines), and a fleet audit found that of 243 enabled scheduled tasks only 41 — 17% — could name an output distinguishing “it ran” from “work landed.”

Five recurring mechanisms are distilled — all-or-nothing validation, writer/reader contract drift, the defaulting accessor, the liveness proxy that measures the guard instead of the guarded, and proof-by-intention — together with the cheap deterministic invariants that turn each from silence into an alarm. Evidence is tiered explicitly: three cases ship with reproducible stdlib-only repros and guards shown red on the broken behavior, two rest on public bug reports, six on the authors' own instruments.

This is an experience report over one fleet (N=1), not a benchmark: no scoring, no vendor comparison, no leaderboard. Companion to the fleet experience report (doi:10.5281/zenodo.22639714).

Code and data: github.com/tonydzi/agent-control-plane-casebook (MIT). Contact: dzyatkovskiy.a2@gmail.com · ORCID 0000-0001-7408-3054 · github.com/tonydzi · tonydzi.github.io.

Keywords

LLM agents · multi-agent systems · control plane · scheduler · task registry · watchdog · silent failure · experience report · reliability · observability · reproducibility · casebook

How to cite

Dziatkovskii, A. (2026). Silent by Construction: A Casebook of Control-Plane Failures in a Production LLM Agent Fleet. Preprint. Zenodo. https://doi.org/10.5281/zenodo.22688465

@misc{dziatkovskii2026controlplane,
  author    = {Dziatkovskii, Anton},
  title     = {Silent by Construction: A Casebook of Control-Plane Failures in a Production LLM Agent Fleet},
  year      = {2026},
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.22688465},
  url       = {https://doi.org/10.5281/zenodo.22688465},
  note      = {Preprint, CC BY 4.0}
}

One text, one DOI: the PDF on this page is the Zenodo file byte for byte. Please cite the DOI.

Anton Dziatkovskii · Palo Alto AI Research Lab · All 2026 preprints · Full publication list