Every hand-written seam needs a guard. Crossweft, a cross-language contract check for AI agents
By Anastasiia Butova, ComfyUI and diffusion-model engineer, Belgrade · published 8 October 2026
Our most expensive bugs were never inside a file. They were between files. One value was written in two or three places, in two languages, someone changed one of them, everything still built, the tests were green, and the install broke on a customer's machine. This is how we made those places explicit and checkable, and built the check into the AI agents that write a large share of our code. The tool is open source: github.com/AnastasiyaW/crossweft, project page: happyin.work/crossweft.
Ten boundaries the compiler can't see
We build a retouching plug-in for Photoshop that runs a neural network on the user's own machine. Between the photographer and the model sit ten trust boundaries: a protected channel between the plug-in and a system service, device enrolment, signed artefacts with trusted timestamps, build attestation, the update channel, model encryption and a watermark. The parts are written in C++ and Go, run in different processes and arrive through different installers.
None of these boundaries is a module. Each one is an agreement between two sides: the same service name, the same protocol version, the same set of routes, the same confirmation string. Today 54 links across those boundaries carry 188 value comparisons (468 places in the code) and 105 set comparisons. The compiler sees none of them, because the two sides are in different languages.
We started calling these places seams: points where two components have to agree exactly.
- the route a client calls and the route the server registers;
- a protocol version header;
- the fields of a DTO in Go and in TypeScript;
- the environment variables a service reads and the ones listed in
.env.example; - a price-rounding rule written twice, in two languages.
Types, tests and linters work well inside one language. They rarely check agreements like these, so those agreements live in people's memory.
The name comes from the weft, the thread that runs across the warp and holds the fabric together.
Three stories
One identifier, three copies. The service name was resolved in three nearly identical functions. On a clean machine the service wasn't installed yet, the lookup failed, an empty machine looked like an old incompatible install, and the installer stopped with exit code 76. We fixed one copy. The other two stayed broken.
Two halves of one dependency. The build took a library from one install directory and its tools from another. The configuration only printed a warning, so the problem surfaced much later, where the real cause was hard to find.
A guard that guarded nothing. A validator for one of our security rules had a wrong path. For weeks it checked zero files and reported OK every time. It passed precisely because it didn't run.
The cause is the same in all three: an agreement has two or three sides, and whether they match depends on someone paying attention.
Why AI agents make this worse, and how they help
A large share of our code is written by agents, mostly Claude Code and Codex. They handle cross-component agreements worse than people do. The typical case: the agent changes a Go struct, runs the tests, sees green and says "done", without ever opening the TypeScript interface on the other side.
- Local edits. The other side of the seam is in another folder, another language, another process.
- Limited context. The agent doesn't hold the whole product in its head, and the compiler won't tell it that a client somewhere depends on that struct.
- A narrow idea of "done". For an agent, done means "every check I can see has passed". If no check looks at the seam, the drift stays invisible until something breaks.
Agents follow precise local instructions well. Tell an agent, right after the edit, which file on the other side to open and what must match there, and it will do it. Don't let it finish while the two sides disagree, and the mistake never reaches a commit.
The rule: every link says how it is kept in agreement
Every link between components declares how its two sides stay in agreement. If both sides are written by hand, the link needs a guard: a check that reads both sides.
Links live in JSON files next to the code. The map has blocks (programs, services, stores, config files) and links between them (HTTPS, a named pipe, files, code generation, environment variables). Every claim in the map is anchored to a literal in a source file, so when code moves, the map can't keep describing the past in silence. The check fails.
Each link has a contract.enforcement field:
| Enforcement | Who keeps the sides equal | What the check requires |
|---|---|---|
shared-code |
a symbol both sides import | defined_in names the symbol |
generated |
a generator from one source | defined_in names the source |
schema-tests |
shared test vectors | defined_in names the vectors |
duplicated |
nobody, it's copied by hand | a guard that reads a file on each side |
convention |
nobody, it's agreed | a guard that reads a file on each side |
none |
nobody | a guard that reads a file on each side |
The classification alone is useful. "Where do we still rely on someone's memory?" becomes a list you can count. A guard that reads only one side doesn't count: the tool checks that it reads files at two different ends of the link.
Three guards: join, set and pair
The guards are deliberately simple: regular expressions, comparing values or sets of values. A check takes milliseconds, needs no build and works with any language.
join: the value is the same everywhere. The API version the TypeScript client sends and the one the Go server expects:
{"id": "api-version", "link": "web-orders", "points": [
{"path": "web/src/api.ts", "side": "web", "regex": "API_VERSION = \"([^\"]+)\""},
{"path": "server/main.go", "side": "api", "regex": "APIVersion = \"([^\"]+)\""}]}
Every occurrence is compared, not just the first. Our first version checked only the first match, so a second copy further down a file could slip through.
set: the members match. TypeScript interface fields against Go JSON tags; the statuses a Python worker writes against the Go constants; the environment variables the code reads against .env.example. An intentional difference goes into allow with a reason, and an allow entry that is no longer needed fails the check itself.
pair: duplicated logic with no single value to compare. For example, price rounding in Go and in TypeScript. Both regions are marked with crossweft:begin price-rounding and crossweft:end price-rounding, and Crossweft fingerprints them. If either changes, the check fails until someone re-reads the twin and runs crossweft attest price-rounding --reason "...". The reason goes into the lock file and into review.
The shortest way out of a red pair for an agent is to attest the side it changed and never touch the other one. So if only one region changed since the last attestation, attest refuses. Fix the twin first, or, if it really is already equivalent, say so explicitly with --other-side-unchanged.
There are bigger checks too. Every route the server registers has to be on the map; Go chi is read natively (also Express and Fastify, FastAPI, Flask, gin and echo), anything else through a regular expression. Every /v1/... call in client code has to be explained by a link. A block can only send data it created or received. Every source folder has to be on the map.
When the built-in kinds aren't enough, a repository adds its own as plug-ins. We have 13. One compares 916 places where the Go server talks to SQL with 161 tables and 276 triggers in the schema, another checks 185 server routes, a third checks 27 "waiter ≥ work + margin" pairs for timeouts. A plug-in that fails to import, crashes or compares nothing is an error, never a green result.
What the agent sees
The check is half of it. The agent needs to hear about a seam right after the edit, not when the code is ready to commit. Three hooks do that:
- Session start: a short note that the repository has a seam map, how many seams, and how to work with it.
- After every edit: if the file sits on a seam, the agent is told the other side and the guard.
- Before finishing:
crossweft checkruns. If a seam disagrees, the agent can't finish and is told what to do.
After an edit to the Go server in the demo, the agent sees:
crossweft: server/main.go is part of 3 cross-component seam(s). Keep both sides in agreement:
- link web-orders (web -> api, https; Orders API [duplicated]): you changed its to side.
Re-read: web/src/api.ts. Guarded by: join:api-version, join:version-header,
set:order-fields, set:web-statuses, pair:price-rounding.
...
Run `crossweft check` before you finish.
If it bumps the API version in Go and tries to stop, the hook sends it back:
crossweft check fails:
- join:api-version disagrees: web@web/src/api.ts:3='2026-09-01'; api@server/main.go:14='2026-10-01'
[server/main.go:14] (key: join:api-version:b5dfb7e5)
A seam is out of agreement: bring the other side in line with the one you changed
(`crossweft impact <file>` names it).
The advice depends on the kind of problem. If a literal the map points at has disappeared, the agent is told to update the anchor or restore the code, not to "fix the other side". The key ends in a digest of which file holds which value, so a recorded exception excuses only that exact disagreement.
A stop hook has two classic failure modes: an endless loop and a silent pass. Ours sends the agent back only while it's making progress (the set of failing keys changes), at most four times in a row. With no progress the agent may stop, and the person gets an explicit message that the check is still red.
For Claude Code there is a plug-in with hooks and a skill. For Codex there is a skill plug-in. Any agent that speaks MCP can connect crossweft mcp and get the check, change impact and "which file to re-read" as read-only tools. The hook adapters for Gemini CLI, Copilot and Cursor are still experimental: their output format is tested, the full loop isn't yet.
Checks that can't pass in silence
The third story was a validator that scanned nothing for weeks. Not every rule is "A equals B"; sometimes you need to prove that a constant has one source. Rules like that live in small scripts, and those are exactly the ones that break quietly.
The crossweft validators runner fails a validator that exits non-zero, doesn't print SCANNED: files=<n> items=<n> or prints zeros, has no self-test or fails it (SELF-TEST: checks=<n>, run against a planted error), or runs out of time. A validator that genuinely doesn't apply prints APPLICABILITY: <reason> and gets SKIP, not PASS.
Crossweft itself works the same way. There are three outcomes, not two: 0 means everything agrees, 1 means drift, 2 means the check didn't happen (the model can't be read, nothing was scanned). An exit code of 2 never looks like success.
A register of known drift that doesn't rot
Not every disagreement can be fixed in the same pull request. Then you record a finding: the problem key, an owner, a next step. While the problem is there, the check passes. Once it's fixed, the finding goes stale and fails the check until someone closes it. Broken anchors can't be excused this way: the map isn't allowed to describe the code wrongly.
For a repository that already has drift, crossweft baseline records it all as open findings with an owner. The check goes green, new drift still fails it, and anything fixed asks to be closed.
What it changed
Honest context first. From mid-July our installer never brought the system service up on a clean virtual machine. We added checked seams on 26–28 September, and on 29 September the service started for the first time. That's a coincidence in time, not proof of cause: other fixes went in at the same time.
What did change is the kind of failure. Three of the first failures on 28 September were seams: exit code 76 (the service name in three copies); the server refusing a loader that registered under the artefact version but was requested under the protocol version; and HTTP 500 because two publishing paths removed the old package and the third didn't. Each now has its own guard. The next two problems were ordinary bugs, not "one value in several places".
What the map shows, recounted from the repository history:
| 27 Sep | 29 Sep | 4 Oct | |
|---|---|---|---|
| Blocks / links | 130 / 160 | 134 / 166 | 138 / 192 |
| join guards | 10 | 220 (529 points) | 261 (644 points) |
| set guards | 4 | 125 | 138 |
| Findings | 58, all open | 106: 93 closed, 10 open, 3 accepted | 146: 117 closed, 25 open, 4 accepted |
The first pass over "every link declares its enforcement" was done by the agents themselves. In one day they added 185 join and 107 set guards and found 15 real disagreements in the code. Of 93 closed findings, 91 have a fixing commit, and 58 of those commit messages name the finding id from the map.
A map without a check goes stale in a week. A week after the first measurement there were 96 problems, and 94 of them were the map going out of date, not bugs in the code. That's why the map is checked every time an agent stops. We went from 4 validators in June to 50, and every one of them has to prove it scanned something.
Opening it up, and what the first CI run caught
The tool grew inside a commercial product, so we opened it carefully. There are two branches: the public oss, which is only the tool, and a private internal, which is oss plus a folder with our own rules. Changes flow from oss to internal only. What goes public is a snapshot of the tree, one commit per release with a neutral signature, not our history. Before publishing, a scanner checks the snapshot and its commit metadata against a list of private words, and a fresh independent reviewer reads the diff since the last release, because a scanner only knows the words it was given. The first review found three phrases in the docs that hinted at how our product is built, and we replaced them before release.
A release tag goes on only after CI is green on Linux, Windows and macOS, because tags are immutable. The very first run was red. It caught two real bugs that 157 local tests had missed:
- the code used an argument that only exists from Python 3.10, while we promise 3.9;
- a plug-in edited twice within one second to the same length ran from stale bytecode, because Python's cache tells files apart by seconds and size. We couldn't see it locally: bytecode writing is switched off machine-wide on our development box.
We fixed both, the second run was green, and v0.1.0 went out. Two days later v0.2.0 added the MCP server, check --changed, guards built from OpenAPI and proto schemas, and route scanners for five more frameworks. It shows the point of this article: a green check only tells you about the kinds of mistake someone has already thought of.
Doesn't this exist already?
Each piece exists somewhere. What we didn't find is the combination, and the rule that every link declares how it is kept in agreement and every hand-written seam has a guard.
| Tool | What it does | How it differs |
|---|---|---|
Google LINT.IfChange, ifttt-lint, if-changed |
"if this block changes, change that file too" | checks that the diff touched the file; a join compares values on every run |
| clevis | one value equal across files | no sets, no regions, no map |
| Pact, buf, oasdiff | contract tests and schema compatibility | the right answer when a schema exists; Crossweft builds guards from OpenAPI and proto for the hand-written side |
| ArchUnit, dependency-cruiser, import-linter | architecture rules | within one language |
| archagent | architecture invariants and skills for agents | closest in spirit, but structure rather than values across languages |
| GitNexus, codegraph | code graphs and change impact for agents, over MCP | a graph you query, not a check that fails; they complement each other |
| fiberplane/drift, Swimm | docs anchored to code | a hash fails on any edit; an anchor fails only when the fact itself is gone |
The approach isn't ours either. Birgitta Böckeler writes about harness engineering: guides that act before the agent writes code and sensors that check afterwards, deterministic checks versus checks done by a language model. OpenAI writes about the same thing in Harness engineering: leveraging Codex in an agent-first world. Crossweft is a deterministic sensor for cross-component seams whose output tells the agent exactly what to fix.
Limitations
- Regular expressions, not parsing. Reformat a declaration and the check fails loudly; you update the pattern. But a pattern that is too broad can still miss a real disagreement, so look at the extracted values and test runtime behaviour with ordinary tests.
- A person still checks that the enforcement is honest. Agents build the map, and Crossweft won't let them invent a link or skip a folder, but whether a link is really
shared-codeor two copies is sometimes something only a developer knows. - Adoption takes work. On a small project the gain is small. It pays off where there are many languages and processes.
- So far we are the only user. The Claude Code loop is proven end to end; other agents connect through MCP and the skill, and their hook adapters are experimental.
check --changedspeeds up pre-commit, but not yet when the edit touches a router or client file.
Try it
The quickest way to see the mechanism is to break the demo (a Go API, a TypeScript client, a Python worker):
pip install "git+https://github.com/AnastasiyaW/[email protected]"
git clone --branch v0.2.3 --depth 1 https://github.com/AnastasiyaW/crossweft
cd crossweft/examples/polyglot-shop && crossweft check # RESULT: PASS
# now change API_VERSION in web/src/api.ts only:
crossweft check # join:api-version disagrees
Each row below was checked on a copy of the demo: the check fails with exit code 1 and names the guard.
| Change | Caught by |
|---|---|
bump API_VERSION in web/src/api.ts only |
join:api-version |
add a JSON field to the Go Order struct only |
set:order-fields |
| add a status to the Python worker only | set:worker-statuses |
read a new env var in the worker that isn't in .env.example |
set:env-worker-vars |
call /v1/refunds from the client |
route:client-unexplained |
change the rounding in server/pricing.go |
pair:price-rounding |
In your own repository: crossweft init (init --example shows a working seam first), then ask your agent to map the repository; the skill walks it through and crossweft check keeps it honest. In Claude Code: /plugin marketplace add AnastasiyaW/crossweft, then /plugin install crossweft@crossweft. In CI: uses: AnastasiyaW/[email protected], or crossweft check --format sarif for GitHub code scanning.
The Crossweft repository guards its own seams with Crossweft: the version in three files, the hook events, the commands in the docs. While we were writing the README, the "documented commands exist" guard fired three times on lines of prose it took for commands. We fixed those too.
Crossweft is Apache-2.0. Issues and pull requests are welcome at github.com/AnastasiyaW/crossweft.