Technical Reportdefuss ·

Verified Agentic Engineering in PracticeAgent-authored control programs and persistent adaptation across a growing component system

Method and data: Aron Homberg. Text: Claude (Opus 5.5), from an editorial specification prepared for the author.

DataThe case-study measurements are frozen at commit 8cecf4b4 (2026-10-07, 17:58 UTC). They cover 33 calendar days, from 2026-09-05 through 2026-10-07. The four engineering changes are measured from the session transcript of 2026-10-06. The general controller implementation is described from defuss-vae 0.8.0 (commit e9b78a3); its later features are distinguished from those used during the study.

Figure 1. The general VAE mechanism. Programs evaluate checks or select bounded work and return instructions to the model through its harness. The agent changes the artifact, invokes the next step and can revise the task-specific programs. A second loop reconciles findings into retained checks, instructions and memory for later sessions. Human policy sets the objective, acceptance criteria and adaptation authority. The diagram describes the method; section 2.3 states the enforcement and initiation boundaries of defuss-vae 0.8.0.

Abstract

Verified Agentic Engineering (VAE) uses agent-authored programs to organize work and return instructions to an agent harness. A verifier names unsatisfied acceptance criteria; a windowed review program selects the next part of an artifact and emits the applicable rules. The shown feedback routines run without a model call. The harness supplies their output to the model, which changes the artifact and invokes the next step under policy established by a human. Findings can also change the method: the agent extends checks, revises instructions and reconciles lessons into reusable tools and bounded prompt memory. VAE-DIALECT records evidence and scope compactly, while stack templates supply established commands and conventions.

This report examines the evolving method in defuss-shadcn, developed by Aron Homberg with diverse model families (Qwen3.8-Flash-Next, Claude Opus 5.5, Claude Fable 5.1) from 2026-09-05 to 2026-10-07. The repository grew from 55 to 231 components across 220 commits and 17 releases. In one session, the agent added typed API documentation for all 59 interactive components (241.9 minutes, one human correction), raised reported line coverage from 31.6% to 75.6% (20.6 minutes), moved 228 component folders under a new layout rule (43.4 minutes), and removed 1,147 type errors (27.3 minutes to final green gates). For the last change, all 85 compared minified artifacts were byte-identical. These observations demonstrate repository-wide adaptation and architectural consistency in agentic engineering under programmed feedback.

1Contribution and scope

Agentic Engineering faces challenges in maintaining consistency, efficiency and correctness when agents autonomously implement change requests in large-scale software repositories - producing potentially inconsistent or incorrect changes that increase the manual review effort by humans. Verified Agentic Engineering (VAE) is a method that addresses these challenges, trading increased time and token spent for higher correctness and reduced human review burden.

A typical challenging task for agentic engineering is a repository-wide requirement change, which typically creates two sub-goals: changing every affected module and establishing that the requirement holds across the whole resulting system now and in the future. VAE lets an agent work on both. It can implement a change and write the programs that identify incomplete work, check the result and direct the next action. The same interface can organize a document review by presenting one section at a time, or a data task by presenting one partition at a time.

The method joins task correction with persistent adaptation through feedback loops. Within a task, program output directs the agent toward a failed condition or the next bounded unit of work. Across tasks, the agent reconciles findings into checks, instructions and selected memory that later sessions inherit. Explicit module contracts give programs something stable to evaluate, and agent harnesses naturally safe split-points to orchestrate work packages in parallel agent swarms. VAE-DIALECT is a superset dialect on top of natural language that expresses formal logic more symbolically and compactly, helping the model to separate between hypothesis, evidence and unknowns more explicitly. Tech-stack templates provide reusable procedures for the toolchains in use, reducing agentic research and debug detours and thus improving overall agentic engineering efficiency.

This report contributes the description of the implemented method and a case study of its operation at project scale. It documents the control interface, the engineering work completed under it and the ways findings by the agent changed it's own subsequent procedures successfully. The study covers one component system (defuss-shadcn), one maintainer, a diverse model family and 33 days. Its results describe the combined system. They do not isolate the effect of the model, harness, verifier, instruction loop, dialect, templates or memory in isolation.

2Programmable feedback and persistent adaptation

VAE operates through a capable language model, an agent harness and programs that the agent itself can create or adapt within the authority granted for the task. The human establishes the objective and governing policy. Programs make parts of that policy executable and turn the current work state into findings or further instructions. The harness returns those outputs to the model, which interprets them, acts and invokes the next step.

The model generates its responses autoregressively; the harness and programs form the outer work-and-review loop. Intermediate corrections can proceed without a new human prompt for each iteration. The model must still be capable of interpreting the evidence, editing the artifact and improving the relevant tools. The case study measures that working combination.

2.1 General feedback protocol

A controller receives the current artifact or task state and produces an actionable result. It may evaluate a predicate, identify the next work unit, supply local context, select applicable instructions or name the next command. The agent responds by changing the artifact (e.g. code or documentation) or task state. When the task exposes a missing capability in the controller (verifier program), the agent can extend that program and validate the change against the governing requirement.

The common interface is program-generated output returned through the agent harness, when the language model e.g. triggers a shell script for building the code, which itself triggers the verifier program and returns it's instructions back to the language model via stdout. The evidence differs by the verifying controller program: an executable predicate verifier can establish its specified condition (e.g. check for consistent symbol existence or naming), while a document window slicer can organize consistent semantic review of documentation prosa - whose judgments remain with the language model.

ControllerProgram operationOutput returned to the harnessModel workEvidence in this report
Code verifierEvaluate executable acceptance criteriaFailing condition, affected location and repair instructionDiagnose and repair the artifact; extend checks when justifiedRepository case study and controlled output capture
Document windowSelect a text part, preceding context and review rulesBounded text, applicable instructions and next commandEvaluate meaning, edit the part and continueSource-inspected controller and controlled output capture
Spreadsheet partitionSelect a row group or sheet region and applicable task rulesPartition, task instructions and next partitionApply the requested transformation and inspect exceptionsArchitectural illustration
Table 1. Three uses of the same feedback loop inducting controller interface method. Code checks and document windows are inspected implementations in this report.

2.2 Program output and the model boundary

An agent and its verifier controller programs talk to each other through the agent harness, so their exchange reads as a conversation. In this repository the coding agent wrote the project verifier and extends it, on a human instruction such as the layout rule of section 4.4, or as a lesson that reconciliation turns into a rule (section 2.4). Figure 2 shows such an exchange with a built-in defuss-vae rule and the document walk on small fixtures, so the program side can be quoted verbatim.

The outputs below were captured from unmodified defuss-vae 0.8.0 functions on controlled fixtures. The verifier example places the temporary probe marker that the hygiene rule searches for on line 2 of component.ts. Ordinary code detects the marker and formats the failure. After the marker is removed, the same check reports success.

The document controller uses a different operation. vae.py prose --walk selects a bounded part of a page and prints its preceding context, static findings and next command. The first part includes the full review catalog; subsequent parts identify the rules. The catalog has 58 rules covering evidence, logic, terms, relevance, structure, language and typography. The verifier program organizes the review, while the model evaluates the text and makes the necessary edits.

Verifier · vae.py gate
VERIFIED[hygiene.probes]=false BC glob='*' files=1; hits=['component.ts:2']
PROVEN: ∅
REMAINS: hygiene.probes
AGENT_CMD: FIX hygiene.probes: component.ts:2
Agent
Removes the probe marker on line 2 of component.ts and calls the gate again.
Verifier · vae.py gate
VERIFIED[hygiene.probes]=true BC glob='*' files=1
PROVEN: hygiene.probes
REMAINS: ∅
AGENT_CMD: ∅
Document review
Document window · vae.py prose --walk
defuss-vae WALK README.md part 2/2 (lines 5-7)
SECTION: # Scope
CONTEXT part 1 (lines 1-4): fix only what crosses into part 2
    1| # Overview
    2| 
    3| The verifier reports failures. The agent applies corrections.
    4| 
REVIEW part 2 (lines 5-7)
    5| # Scope
    6| 
    7| Every required check must pass.
STATIC: ∅
DO: edit README.md in place where a rule fails and the page or repo supports the fix; a fix that needs an absent fact, source or decision is a question for the human, never an invention. Keep structure and voice.
NEXT: python3 vae.py prose --repo . --walk README.md --part 3
Agent
Reviews lines 5-7 against the rules, finds nothing to change and runs the NEXT command.
Figure 2. The agent and its control programs as a conversation. The verifier and document-window messages are captured output from small fixtures using defuss-vae 0.8.0, commit e9b78a3, quoted verbatim; the walk excerpt omits the long CHECK EVERY rule index line between STATIC and DO. The agent's messages state the action each output asks for; they are not captured text. The shown predicate evaluation, window selection and output formatting make no model call. In harness-driven operation, the harness supplies the output to the model's next invocation, where interpretation and the requested editing occur. These captures demonstrate the program interface, not the original production session.

The project's own verifier follows the same pattern with its own wire format: each failing check prints a fix: line naming the command or file to change. When a page figure disagreed with its measurement on 2026-10-07, the stat figures check named bun scripts/stat-figures.ts and a documentation rebuild, and running them resolved the failure.

Its terminal VERIFIED[walk] output depends on the requested terminal part and static findings. The doc and doc-edit skills require following the review procedure; the gate does not validate traversal. A review attestation likewise records the agent's claim, rather than independently establishing the quality of its semantic judgments.

2.3 Harness enforcement and adaptation authority

defuss-vae implements the VAE method as skills, hooks and Python programs. At session start a hook supplies rules, selected stack defaults, repository memory and open leads. The gate runs verification and checks review and documentation attestations for the relevant work state. The repository configures its own commands and deterministic rules in .agents/VERIFY.py, on top of the 87 custom static verification checks implemented in scripts/verify.ts, which VERIFY.py runs as one of its commands. Missing or indeterminate required evidence blocks the monitored commit action.

The harness determines which actions a failed check can prevent. The implementation enforces a commit boundary and supplies continuation instructions. There is no guarantee, that every task runs uninterrupted to completion, but VAE guarantees that every task completion implies a green gate - or in other words: None of the programmatic verifications failed anymore.

Authority to adapt is also policy. The agent may add checks, revise its own procedures and create task-specific tools within the current request. Acceptance criteria still derive from that request and the governing project policy. A passing result obtained by weakening the criterion would not establish the original requirement (no "paperclip optimization"). Eventually, the VAE method guarantees that when a person starts the wrap skill, the agent also carries out its reconciliation procedure - clearing agent memory entries that aren't relevant anymore. VAE also allows for many intermediate requirement change iterations without requiring a fresh intervention at each one.

2.4 Reconciliation and bounded prompt memory

Failures and review findings become episodes with supporting evidence. During reconciliation, the agent compares those episodes with the current code, tests and retained rules. A durable lesson becomes an executable test or verifier rule where that form fits. Otherwise it remains a scoped instruction, memory entry or working command. Later sessions inherit this state and can apply the retained procedure without reconstructing the original failure.

Here, self-learning means changes to the working system's stored evidence, tools and procedures. It does not involve updating model weights. The agent can improve the artifact and the programs that guide or evaluate later work. A retained change needs evidence for its scope. Any current human instruction outranks an older memory entry.

The injected memory is bounded. In this repository the header of .agents/MEMORY.md sets a budget of 4 KiB and that of the command gist 2 KiB; defuss-vae 0.8.0 adds a limit of 240 characters per memory line, and vae.py doctor checks all three. Unresolved leads are retained for reconciliation. When an executable check supersedes a narrative lesson, that narrative can be removed with a reference to the check. Learnings about issues can therefore move from repeatedly injected prose into verifier programs, tools and tests. The total repository of checks and procedures may grow even while prompt memory remains small. In this repository, MEMORY.md held eight entries in 2,311 bytes from commit 7253195d (October 6) through the end of the study.

Agentic memory (retention) alone does not establish improvement in efficiency and correctness; the relevant question is whether the injected memory addresses the previously observed failure mode without weakening the required behavior.

2.5 VAE Dialect and stack templates

VAE-DIALECT supplies a compact symbolic superset dialect over natural language: VERIFIED[scope] identifies a directly supported statement, HYPOTHESIS[scope] identifies an inference to test, and UNKNOWN[scope] identifies a material gap in the available evidence. Operators express conditions, precedence and reasons. The scope and evidence remain part of the statement; confident wording does not turn an inference into an observation. Correctness takes precedence over compression.

These forms are intended to reduce unsupported assertions and ambiguous instructions by forcing the langage model to declare epistemic status explicit. They also reduce repeated phrasing and thus reduce input and output tokens. A VERIFIED label is useful only within the boundary of the evidence behind it.

Stack templates provide concrete toolchain procedures before the agent begins work. The verifier controller program selects defaults for the repository's detected stacks and inserts them into session context or the managed instruction block. When a required command is missing, the verifier includes the corresponding default in its repair instruction. For a Python project, the selected defaults include:

python lint: `uv run ruff check .`
python test: `uv run pytest`
python coverage: `uv run pytest --cov`

Selection and formatting require no model call. These strings supply guidance; the gate executes the repository's configured commands. Templates are intended to reduce repeated tool selection, setup exploration and unnecessary specification work. Along with the dialect and retained procedures, they offer a way to spend more of the context on the current task. However, this report does not isolate their effects on unsupported claims, task errors, token use or elapsed time.

The defaults come as flavors, one per stack: Go, Rust, the JVM, .NET, JavaScript and TypeScript, Python and the web. The author chose them from engineering experience and from research into the preferences of senior developers. They are defaults, not mandates. A repository keeps its existing toolchain unless a person approves a migration. .agents/VERIFY.py can point the lint, test, coverage and end-to-end checks at other commands.

The templates also carry the working procedure. They instruct the skills for planning, implementation, review and documentation, for coordinating sub-agents and for organizing long-running work, with resilience and provenance as design goals: a long job runs detached with a pid file and a log that stamps every line with an ISO-8601 time, the gate records every failure and finding with its session, and every attestation names the code fingerprint it covers.

vae.py init scaffolds a project layout adapted from the Linux Standard Base (LSB) conventions: Makefile verbs (setup, start, stop, status, log, metrics, bench, test, coverage, lint, e2e, verify) with LSB status codes, services that log to var/log/<service>.stdout and .stderr and keep their pid in tmp/<service>.pid, programs that read input/ and write output/, and runtime state kept apart from the code. Sub-agents work in their own git worktrees and share one registry, .agents/SWARM_STATUS.yaml. Every write to it is locked, replaced atomically and read back; only the owning process may change its entry; a claim that overlaps another agent's paths is refused; and liveness is read from the process table, not from the file. The agents can therefore organize themselves without races over the registry and without two agents holding the same paths, a split-brain state.

3Case study and evidence

defuss-shadcn is a component system built from semantic HTML, CSS and JavaScript, with TypeScript sources and no framework or build step required by consumers. It was forked from shadcn-html with 55 component folders. This study follows its development from September 5 to October 7, 2026, under one maintainer and one model family.

Components have explicit contracts. Every interactive component exposes named states, a store, a typed API and a render() operation checked against its authored markup. Shared behavior has an identified owner, including DOM access and component-directory discovery. These contracts let a check range across components and let a shared change reach their callers. The four tasks below exercise that structure at repository scale.

The project had an instruction file and its own verifier from the first commit. defuss-vae was first published on October 1 and wrapped the project verifier from October 6.

The session ran in Claude Code. Every model response recorded inside the four task windows of section 4 came from claude-opus-5-5; later sessions on October 7 also used claude-fable-5-1, a model of the same family.

In Spetember, Qwen3.8-Flash-Next was used to scale from 55 to 77 components.

4Results

The strongest evidence is a sequence of four refactorings in the October 6 session. Each applied a requirement across the relevant component population: typed API documentation, shared-contract tests, the repository's component layout and type correctness. The agent also changed verification for long-term stability under the new requirements and adapted supporting tools to evaluate the new state.

TaskReported endpointMinutesOutput tokensTool callsHuman corrections
Typed API documentationFinal green gate, all 59 interactive components241.9347,6144371
Test coverageFinal reported coverage measurement20.639,926480
Repository layoutFinal green gate, all 228 component folders43.476,358980
Type errorsZero diagnostics and final green gates27.382,8551300
Table 2. Four changes from one session, measured from each request to its reported endpoint. API documentation, layout and type correction end at final green gates; coverage ends at its final reported measurement. The windows include contemporaneous work: section bundles in the first, and a CI workflow and report edit in the second. Times and token counts describe those task windows, rather than isolated controlled trials. Human corrections count correction prompts within each window, not all human input or oversight.

4.1 Growth and verification coverage

Components

from 5505.09.to 23107.10.

E2E test files

from 505.09.to 25207.10.

Unit tests

from 505.09.to 59907.10.

Verifier checks

from 3305.09.to 8707.10.

Commits

from 1005.09.to 22007.10.

Releases

from 005.09.to 1707.10.

Lines added

from 56,64005.09.to 350,33107.10.

Lines removed

from 1,26705.09.to 80,99907.10.

The repository grew from 55 components at the fork to 231 at commit 8cecf4b4. The boxes compare the last commit of September 5 with the cutoff commit; the line counts are cumulative from the fork, whose import they include. At the end of the first day the repository held 5 E2E test files and one unit-test file with 5 tests; at the study endpoint, 252 E2E files and 35 unit-test files, whose last full recorded run held 599 tests. Over the same span the verifier grew from 33 checks to 87, counted as the distinct labels of the check(...) calls in scripts/verify.ts. Its recorded October 5 run, on an earlier tree, passed 78 gate instances, the result lines the verifier prints; warnings are not counted as passes.

The interval contains 220 commits and 17 releases. The largest component-count increase followed the addition of website blocks and application scaffolds on October 2. These counts establish the size and activity of the project during the study; they do not independently measure correctness or the method's causal effect.

Figure 3. Components and E2E test files at the last recorded commit of each calendar week, read from git objects. The first and last intervals are partial. The fixed endpoint is October 7, 2026: 231 components and 252 E2E files. The fifth interval includes the website blocks and application scaffolds added on October 2. The weekly grouping uses the original measurement's UTC+2 convention; task timestamps elsewhere are UTC.

4.2 Refactoring: Introduction of a new, typed API documentation

The first requirement was adding a typed API documentation for all 59 interactive components. Every member needed a TypeScript signature, every argument a description, every event typed detail, and every state its accepted configuration. The shared State API also had to be typed for each component's own states and checked in both source and rendered documentation.

The final green gate was reached after 241.9 minutes (including 1h computer hibernate caused by low battery), 347,614 output tokens and 437 tool calls, with one human correction. All 59 interactive component sources changed. The source verifier checked all 59 interactive components, and the E2E audit opened all 59 rendered API pages. The first green result had missed configuration information; section 5.1 describes how that finding expanded the verification scope.

  1. The requirement arrives

    ...during another task; both get bundled and processed together.

  2. First green gate

    Verify, review and documentation attestations.

  3. One human correction

    The maintainer opens a second page: the configuration each state accepts is missing. The human changes the requirement accordingly.

  4. Final green gate

    All 59 interactive components complete, with an E2E audit of all 59 rendered API pages.

Figure 4. API-documentation timeline on October 6, 2026, read from the session transcript. The first green gate was followed by one human correction; the final gate included an audit of all 59 rendered API pages.

The change added 3,455 and removed 408 lines across the component sources. Across all authored files in the window, 246 files changed, with 7,570 additions and 1,537 deletions; that larger total includes section-bundle work and a documentation pass. The compiler ratchet fell from 1,212 diagnostics to 1,147. These are separate scopes: the documentation requirement covered the 59 interactive components, not every component in the project.

4.3 Task: Increase test coverage

The next task raised reported browser-mode Vitest/V8 line coverage from 31.6% to 75.6% in 20.6 minutes, using 39,926 output tokens and 48 tool calls. Recorded covered lines increased from 5,460 to 15,880. The work added 147 lines of tests and an 11-line component fix. The measurements include loaded documentation code and third-party packages; the percentages are those of their respective instrumented runs.

The agent reused 63 existing interactive fixtures: the 59 component fixtures and four compositions. The first contract test exercised every declared state and checked agreement among getState(), the store and render(), while requiring an undeclared state to throw. It exposed a missing settled() operation in the cookie-consent controller. The second test operated controls and fields and required that handlers not throw and promises not reject unhandled.

MeasureBeforeAfter
Lines, step 1: states31.6%up 66.9%
Lines, step 2: controls66.9%up 75.6%
Statements, step 1: states-61.0%
Statements, step 2: controls61.0%up 71.8%
Functions, step 1: states-60.0%
Functions, step 2: controls60.0%up 71.1%
Branches, step 1: states-38.6%
Branches, step 2: controls38.6%up 50.5%
Unit tests, step 1: states420up 497
Unit tests, step 2: controls497up 564
Coverage floor the gate enforces, step 1: states30%up 60%
Coverage floor the gate enforces, step 2: controls60%up 70%
Table 3. Coverage measured with Vitest/V8 in browser mode before the task and after each of its two steps, each row comparing two measurements; step 1 was measured at 13:12:56 UTC and step 2 at 13:29:08 UTC. Line coverage rose from 31.6% to 66.9% and then 75.6%; the enforced floor rose from 30% to 60% and then 70%. The unit-test count rose from 420 to 497 and then 564. Unrecorded baseline metrics are marked "-". Each percentage belongs to its recorded instrumented population.

The shared fixtures and contracts made it possible for two tests to exercise all interactive components. The assertions give evidence about the stated contracts and error paths; the line-coverage percentage records the proportion of instrumented lines executed. Other behaviors and paths remain outside those assertions.

4.4 Refactoring: Structural change of the repository layout

The third request was to make the source tree follow the documentation sidebar. Every component needed a section: entry in its skill, a matching section directory and a stable published path. The agent moved all 228 component folders into 19 sections in 43.4 minutes, using 76,358 output tokens and 98 tool calls, without a human correction.

The change introduced a shared owner for directory discovery and rewired 11 tooling readers to use it. It affected 611 authored files, with 1,008 additions and 625 deletions after rename detection. The compiler ratchet remained at 1,147 diagnostics; 570 unit tests and 246 E2E files passed at completion. Intermediate failures in path discovery, generated links and configuration were corrected before the final quality gate.

MeasureCount beforeCount after
Section folders under src/components/0up 19
Component folders inside a section folder (moved with git mv, history kept)0up 228
Files in the component folders (all moved, none added or removed)582582
Skills declaring section:0up 228
Sibling links routed through the section0up 261
Component sources importing from the section depth0up 59
Tooling readers using the shared directory owner (component-dirs.ts)0up 11
Documentation components rewired-4
Tests rewired-8
New modules-3
New test files-1
Authored files changed (renames detected)-611
Lines added-+1,008
Lines removed-−625
Type diagnostics (the compiler ratchet)1,1471,147
Table 4. The repository-layout change, from git and verifier output. All 228 folders, containing 582 files, moved under 19 sections. Published component paths remained flat. The recorded minified sizes were unchanged; this measurement alone does not establish byte identity. An arrow marks the direction of a change; "-" marks a count that has no before value.
src/components/
  button/
    button.css
    component-skill.md
  data-grid/
    component-skill.md
    data-grid.css
    data-grid.schema.json
    data-grid.ts
  dialog/
    component-skill.md
    dialog.css
    dialog.schema.json
    dialog.ts
  ... 225 more component folders
Before
src/components/
  actions/button/
    button.css
    component-skill.md
  big-data/data-grid/
    component-skill.md
    data-grid.css
    data-grid.schema.json
    data-grid.ts
  overlays/dialog/
    component-skill.md
    dialog.css
    dialog.schema.json
    dialog.ts
  ... 225 more, in 19 sections
After
Figure 5. Three of the 228 component folders before and after the move (commit cf69c1b5): each folder keeps its files and moves into the section that lists it in the sidebar. Drag the divider to see more of either tree.

The new layout rule has three negative cases: a missing section:, a component in the wrong section and a component outside all sections. Each fails with an instruction identifying the move or front-matter correction. This result shows the agent changing the repository and making the new structural requirement executable for later changes.

4.5 Task: Fix type errors under output equivalence

Later in the same session, the agent was asked to remove the remaining 1,147 compiler diagnostics without changing the shipped output. It reached zero diagnostics in 12.5 minutes, using 71,958 output tokens and 100 tool calls. The final green gates were reached after 27.3 minutes, 82,855 output tokens and 130 calls, with no human correction.

A type change in the shared query module removed 707 diagnostics before any component was edited; the other 440 required local corrections. The complete change touched 45 component sources and three shared or type files, adding 276 lines and removing 210. At completion, 571 unit tests and 246 E2E files passed, and all 85 compared minified artifacts were byte-identical to the preceding build.

MeasureCount beforeCount after
Type diagnostics in the component program1,147achieved 0
... after the type fix in the shared query module1,147achieved 440
... after the local corrections440achieved 0
Component files with diagnostics52achieved 0
Component sources changed-45
Shared modules and type declarations changed-3
Source lines added-+276
Source lines removed-−210
Source-reading tools that accept a type argument0up 3
Shipped minified artifacts, byte-identical to the previous build85achieved 85
Table 5. Type-correction results from git, compiler output and byte comparison. Equality applies to the 85 minified artifacts compared before and after this change. The readable JavaScript differed in two lines because an erased cast left parentheses that the minifier removed. A check marks a result the change aimed for: fewer diagnostics, identical bytes. An arrow marks the direction of any other change; "-" marks a count that has no before value.

4.6 Human review effort

On September 7 the maintainer reviewed the documentation site for about eight hours and filed 52 issues, including missing examples, field hints, dead links and visual concerns. Eighteen issues were closed within twelve hours. The remainder were closed in a batch on September 28, so those closure timestamps are upper bounds on the corresponding fix dates.

4.7 Changes to committed code (Code Churn Rate)

Across the study interval, git recorded 350,331 additions and 80,999 deletions in authored files, excluding generated output. The first day's 56,640 additions include the imported fork. Documentation accounts for 69,338 deletions, including the migration from static pages to MDX.

Across components, tests, the shared runtime and theme, git recorded 8,079 deleted lines and 165,668 added lines, a deletion-to-addition ratio of 4.9%. These aggregate counts do not measure the survival of individual lines. They describe committed change volume, including migrations, rather than a defect or rework rate.

AreaAddedDeletedNetDeleted / added
Components (src/components)+99,100−5,793+93,3075.8%
Documentation (src/documentation)+166,874−69,338+97,53641.6%
Tests (tests/)+57,578−2,029+55,5493.5%
Tooling and verifier (scripts/)+10,726−1,910+8,81617.8%
Theme and tokens (src/theme)+5,786−119+5,6672.1%
Shared runtime (src/shared, src/core, src/types)+3,204−138+3,0664.3%
Other (instructions, skill templates, configuration)+7,063−1,672+5,39123.7%
All authored files+350,331−80,999+269,33223.1%
Table 6. Added and deleted line events by area, from git log --numstat over September 5 through October 7, with rename detection and generated files excluded. The initial fork import is included. The ratio is deletions divided by additions.

4.8 Execution and context cost

The recorded October 5 run of the project verifier evaluated 78 gate instances in 10.2 seconds on an Apple M4 MacBook Air. That is the verifier command's duration, not the complete build, screenshot and test pipeline. Task-window elapsed times and output-token counts are reported separately in Table 2.

The session that produced the first version of this report also records input-token accounting. Across its requests, 580,660,897 input tokens were cache reads, 2,609,778 were cache writes and 2,330 were uncached input. Cache reads account for 99.6% of the sum. These are accumulated usage counters, including repeated reads of cached context; they are not the size of a single context or a count of unique text.

The three categories are priced very differently. At the Claude API list prices for Opus 5.5 at the time of writing (Anthropic, 2026), a cache read costs $0.20 per million tokens, 95% below uncached input at $4 per million. A cache write costs more than uncached input: $5 per million for the default 5-minute cache and $8 for the 1-hour cache, which every cached prefix of this session used. At these prices the recorded input comes to $137.02: cache reads were 99.6% of the tokens and 84.8% of the cost, cache writes 0.4% of the tokens and 15.2% of the cost. The same tokens at the uncached price would cost $2,333.09. These are list-price figures for the recorded counters, not the amount the session was billed.

Input usage categoryPrice per millionTokensCost
Cache reads: least expensive, 95% below uncached input$0.20580,660,897$116.13
Cache writes, 1-hour cache: most expensive, twice uncached input$8.002,609,778$20.88
Uncached input: the base price$4.002,330$0.01
All input-583,273,005$137.02
Table 7. Input-token accounting for the session that wrote the first report version, through the recorded October 7 measurement (04:20 UTC), priced at the Claude API list prices for Opus 5.5 read on October 8, 2026 (US dollars). Every cached prefix of the session went to the 1-hour cache. The first request recorded 69,980 cache-write tokens. These counters include the session context and cannot be attributed exclusively to the instruction file. The costs are list prices applied to the counters, not the amount billed. These prices apply to token-based (API) billing only. A monthly subscription charges a flat fee within usage limits instead, so for an individual developer on such a plan the per-token prices matter only through those limits. The maintainer reports working on the Claude Max plan of 2026 and, across this and several other projects, never exceeding 90% of its weekly limit.

Caching, compact instructions and reusable procedures address different costs. The observed cache counters describe how input was served. They do not measure the independent token savings of VAE-DIALECT or stack templates, and this report does not calculate a monetary saving from them. A reasonable hypothesis follows from their mechanisms: example templates that work with the project's tooling spare the agent derailing loops of failed attempts, and symbols in place of long-form words shorten every instruction, so each should reduce token use on its own. The mechanisms support that isolated reduction, but no measurement does yet, and the combined effect of all parts of the VAE method has not been measured; the comparisons of section 6 would test both.

5How project/task-specific self-learning adapts the method and makes it scalable

The following cases distinguish three responses to a finding: extending an executable check, retaining a behavioral instruction and obtaining a human acceptance decision. They are selected to explain the method, rather than to estimate an error distribution. A discovered failure contributes to persistent improvement when its correction is validated and retained at the appropriate scope.

5.1 Extending the verifier

The typed API documentation first passed its gate at 06:59 on October 6. At 07:02 the maintainer opened a second page and found that the configuration accepted by each state was missing. The check had covered source information without establishing that every rendered page exposed it.

Under the VAE method, the agent corrected the documentation and extended the work to an E2E audit of all 59 interactive-component pages (tests/e2e/api-docs.e2e.ts, added in commit 7253195d). The final green gate was recorded at 09:29. The human supplied the missing acceptance observation; the agent supplied the repair and the broader check. The correction therefore changed both the artifact and the verification scope. Once an executable check embodies that requirement, later work can test it without reconstructing this exchange.

5.2 Encoding ownership in persistent instructions

The project recorded failures at module boundaries: the dialog script claimed a dialog owned by another component, and the tabs script claimed the BibTeX component's tab list. The first version of this report analyzed them under the heading "Small modules, clear contracts". At 03:56 UTC on October 7 the maintainer asked to make that guideline a strong recommendation in AGENTS.md; the agent drafted it, the maintainer shortened the wording at 04:19, and commit 10d02a3c added it at 04:38. The instruction states that one module owns each concern and other modules use that owner. The same rule entered the plan and implement skills of defuss-vae 0.6.0 at 08:49 that day.

The record therefore links the observed failures, through the report's analysis and a human request, to the persistent instruction. In a later session the agent placed a new README-parity rule in the module that already owned the README parity checks (commit a658009a), as the rule directs. That is one recorded application; the record does not measure how far the rule reduces ownership errors. The runtime repairs remain separate from the persistent instruction change. The same skill update also carried two lessons from section 4.5 into implement: compare built artifacts when a change must not alter behavior, and test every source-reading tool when source syntax changes.

5.3 Human acceptance

Some findings require a preference or acceptance decision that the current specification does not contain. A verifier can check a stated radius, alignment constraint or timing requirement once it is encoded. It cannot derive the author's intended appearance merely from a successful structural check. The human must supply that preference or authorize a source of judgment; the agent can then apply the decision and encode the parts that admit a stable check.

One such decision concerned this report's own typesetting. Its sections set the opening line of each first paragraph in small capitals, a Typography option that applied as specified while the paper and typography tests passed. On October 7 the author judged the result ugly: the line read as monospace. The agent traced the appearance to the paper's font, Geist, whose shipped file declares neither of the OpenType small-capital features (smcp, c2sc), so the browser synthesized them from shrunken, letter-spaced capitals. The author's criterion was appearance, not a failed requirement. The agent removed the option from the report's sections, and the change shipped in commit 70628bc9. No check was added. CSS cannot tell whether a font provides real small capitals; a check over the font's feature tables would be possible, but none exists in the project.

The visual record of October 7 marks the same boundary from the other side. Four screen-reported defects became E2E checks that day. A rotating-glyph artifact remained unresolved because its edges had been measured only in a still frame.

6Scope and remaining limits

The results presented in this paper come from one project, one maintainer and one model family, without a control group. The working method changed during the study, and the general package (defuss-vae) was introduced near its end as a generalization of the VAE method engineered within the defuss-shadcn project over time. The task windows also include contemporaneous work and are not controlled latency or cost benchmarks.

Checks establish their encoded predicates over their measured inputs. A green result does not establish unspecified behavior, and a review attestation does not independently establish review quality. These boundaries motivate extending coverage and improving instructions when a new failure is understood. Changes to the verifier itself must preserve the governing requirement; making a check easier to pass is not evidence that the original task was solved.

The control interface is independent of a particular artifact format: programs can return repairs, document windows or data partitions through the same harness interface. This report demonstrates code-related work and inspects the document controller. Transfer to spreadsheet tasks, other repositories, teams, languages or model configurations remains an empirical question.

In scope of future research, the next comparisons should hold the task, model, tool access and acceptance criteria fixed while varying the feedback controller, persistent adaptation, dialect and stack templates. These comparisons would identify which parts of the combination contribute to the observed performance most.

7Self-learning and agent memory

Memory is where self-learning either compounds or decays. Two problems grow with a project. A memory searched at run time has to retrieve the right entries from a store that keeps growing. A memory injected whole keeps entries that have gone stale, and a stale entry keeps steering the agent after the code has moved on. Ning et al. (MemAdapter, 2026) describe a related effect: persistent memories can make an agent over-align with a user's past beliefs, even when those are inaccurate, outdated or contradicted by evidence.

VAE splits memory in two. Episodic memory is .agents/EPISODES.md: the gate appends a timestamped FAIL, DONE or FINDING line, with its session, for every outcome. At most three of its entries are injected: the newest open ones, which are unsettled lessons and findings, and failures that no later success of the same session resolved. Agents search the rest by path or symptom when a task touches the same place. Long-term memory is small and injected whole into every session: .agents/MEMORY.md, one line per decision with its reason and scope, and the command gist, within the budgets of section 2.4. Most of what the project has learned sits outside the prompt altogether, in tests and verifier rules that every later gate runs.

Every session started in the repository receives this state at start, the orchestrator's and each spawned sub-agent's, and every gate run writes episodes. Reconciliation happens in wrap, which a person starts when a work package is done; spawned sub-agents never run it. It reflects first: each lesson the work taught and each open lead is promoted to the strongest form that fits, a test, a verifier rule or one memory line with its reason, or deleted with evidence. It then audits every entry against the current code and commands. A contradicted entry is rewritten or deleted, an entry a test now enforces is deleted, one wider than its evidence is narrowed, and one it cannot settle stays, tagged UNKNOWN. The aim is that what reaches the next session is still true, still useful and small enough to load in every session; for the last part, wrap runs vae.py doctor, which checks the length and count limits. Two lessons from the days after the study took the strongest form, a test: on October 7 a deadlock between two gates became four unit-test cases, and on October 8 a Firefox scroll loop became an end-to-end test.

Ning et al. also find that a correct memory can still induce sycophancy, and that the same memory warrants different influence in different contexts. MemAdapter answers with three reasoning components that weigh each retrieved memory against the current task and the evidence. defuss-vae answers in the rules for using memory: an entry binds only within the scope its evidence supports, the current request outranks it, and repetition, recency or confident wording widen nothing.

The approach has limits. Reconciliation is only as good as the agent's judgment during wrap, and its attestation records a claim, not its quality. It runs when a person starts it, so lessons wait for the end of a work package. A lesson that cannot be stated as a check stays prose, and the budget forces a choice among such lessons. This study does not measure how much stale or wrong memory the audit removes, or whether retained lessons prevent repeat failures; section 6 names the comparison that would in scope of future research.

8Conclusion

Over 33 days, defuss-shadcn grew from 55 to 231 components while its verifier and instructions evolved with the project. In one session, the agent applied typed API documentation to all 59 interactive components, expanded shared-contract tests, moved all 228 component folders into a new layout and removed 1,147 type errors. For the type change, all 85 compared minified artifacts remained byte-identical. One human correction to the documentation task became a broader audit of the rendered pages.

The contribution of this work is the demonstration of the working combination: a capable language model for coding in an agent harness, agent-authored programs that return actionable instructions, explicit contracts, compact evidence notation, stack templates and reconciled memory. The agent can change the task artifact and, within its authority, the tools and procedures that guide later work. Lessons learned during execution can become executable checks or persistent instructions, while agent memory reconciling helps to prevent stale records, allowing knowledge to remain in the system without retaining every episode in the prompt.

This use case demonstrates substantial repository-wide work under the Verified Agentic Engineering (VAE) method. Its general interface supports further tasks that can be expressed as checks or bounded instructions, stressing that this method is - in principle - not limited to agentic engineering alone.

Citation and references

Cite this report
@techreport{homberg2026vaepractice,
  title = {Verified Agentic Engineering in Practice: Agent-Authored Control Programs and Persistent Adaptation Across a Growing Component System},
  author = {Homberg, Aron},
  institution = {defuss},
  year = {2026},
  month = oct,
  note = {Text drafted with Claude (Opus 5.5)},
  url = {https://kyr0.github.io/defuss-shadcn/paper.html}
}
References
@misc{homberg2026defussvae,
  author = {Homberg, Aron},
  title = {defuss-vae: Verified Agentic Engineering},
  year = {2026},
  version = {0.8.0},
  howpublished = {GitHub},
  note = {Commit e9b78a321a70446b4dd7d00bf84f65b162d76615},
  url = {https://github.com/kyr0/defuss-vae/tree/e9b78a321a70446b4dd7d00bf84f65b162d76615}
}

@misc{homberg2026defussshadcn,
  author = {Homberg, Aron},
  title = {defuss-shadcn: the study repository},
  year = {2026},
  howpublished = {GitHub},
  note = {Study cutoff at commit 8cecf4b4, 2026-10-07},
  url = {https://github.com/kyr0/defuss-shadcn/tree/8cecf4b4}
}

@misc{ning2026memadapter,
  author = {Ning, Ruqing and Meng, Haibo and Xiang, Zhishang and Chen, Zerui and Su, Jinsong and Wang, Xin and Zhang, Qinggang},
  title = {MemAdapter: Counterfactual Adaptation Against Memory-induced Sycophancy},
  year = {2026},
  eprint = {2610.05162},
  archiveprefix = {arXiv},
  doi = {10.48550/arXiv.2610.05162},
  url = {https://arxiv.org/abs/2610.05162}
}

@misc{anthropic2026pricing,
  author = {{Anthropic}},
  title = {Pricing - Claude Platform Docs},
  year = {2026},
  url = {https://platform.claude.com/docs/en/about-claude/pricing},
  note = {Accessed 2026-10-08}
}

@misc{lindley2026shadcnhtml,
  author = {Lindley, Cody},
  title = {shadcn-html: The AI prototyping substrate},
  year = {2026},
  howpublished = {GitHub},
  url = {https://github.com/codylindley/shadcn-html}
}

@inproceedings{park2021nerfies,
  author = {Keunhong Park and Utkarsh Sinha and Jonathan T. Barron and Sofien Bouaziz and Dan B. Goldman and Steven M. Seitz and Ricardo Martin-Brualla},
  title = {Nerfies: Deformable Neural Radiance Fields},
  booktitle = {Proceedings of the IEEE/CVF International Conference on Computer Vision},
  pages = {5865--5874},
  year = {2021}
}