Use case

Fix flaky tests with Codna

A retry hides a flaky test. Codna finds the shared state or ordering it depends on, fixes the cause, and re-runs your tests until they pass.

The problem

A retry is not a fix.

Flaky tests fail because of something outside the test: shared fixtures, ordering, timing. Codna builds a dependency and blast-radius graph of the repository from its import patterns. No model call, zero tokens. The agent starts from the failing test and the code it really depends on.

How Codna fixes it

How Codna fixes it

1

Map the suite

Codna builds a dependency and blast-radius graph of the repository from its import patterns. No model call, zero tokens.

2

Fix the cause

The agent receives an evidence bundle scoped to the issue: the suspect files, the call paths, the failing test. Codna prints the raw-to-bundle token size on every run. Every fix reports root cause, confidence, blast radius and regression risk, and passes a risk gate before it is applied or a pull request opens.

3

Re-run until green

Run codna fix --tests --apply and Codna runs your tests in a sandbox and re-fixes until they pass, up to the iteration limit you set. Set fix.test_command in codna.yaml, or pass --test-cmd, when your runner is not pytest. With --open-pr, or through the GitHub App, the pull request states the issue, the root cause, the symbols touched and a confidence score, and asks for review before merging. Codna never merges.

codna fix . --issue "test_orders is flaky under parallel run" --tests --apply --max-iterations 3

What you get

What you get

The cause, not a retry

The agent works from the failing test, the call chain and the suspect fixtures, so the patch targets shared state or timing.

Tests in the loop

Run codna fix --tests --apply and Codna runs your tests in a sandbox and re-fixes until they pass, up to the iteration limit you set. Set fix.test_command in codna.yaml, or pass --test-cmd, when your runner is not pytest.

Blast radius reported

Codna reports what the change reaches before you apply it.

The proof

Fewer tokens. Faster. Verified.

Codna16K
Cursor81K
Average tokens per fix on 87 matched bug-fix cases: Codna and Cursor.

Frequently asked

Codna builds a dependency and blast-radius graph of the repository from its import patterns. No model call, zero tokens. The agent receives an evidence bundle scoped to the issue: the suspect files, the call paths, the failing test. Codna prints the raw-to-bundle token size on every run. The agent fixes the dependency the test has on shared state or ordering.

No. The graph names what the test depends on. --tests runs the suite in a sandbox to discover failures and to verify the patch.

Yes. thyn-ai/codna-action@v1 runs codna fix in your CI and opens a pull request. The GitHub App triages a red check suite and opens a fix on the pull request branch.

On 87 matched bug-fix cases against Cursor, Codna averaged 16,159 total tokens and about $0.02 of model spend per verified fix, 5× fewer tokens and 1.7× faster. With your own key you pay your provider directly.

The repository map is built from import patterns, so it is not tied to one language. On-device recall covers Python, JavaScript, TypeScript, Go, Rust, Java, C, C++, C#, PHP and Ruby. Your own test command verifies the fix.

Understanding runs on your machine and spends no tokens. Only the evidence bundle or the diff reaches your model provider, with your key from the OS keychain. Set privacy.egress to fail-closed in codna.yaml and Codna runs your tests only under kernel-level network denial. Secret redaction is always on.

Understand. Fix. Evolve.