- Similar work is performed again for each task.
- Capacity grows mainly by adding human time.
- Process knowledge may remain informal.
Real tasks in. Verified results out.
FaultWright is being built to autonomously execute permitted software-engineering AI-training tasks and independently verify the result.
Existing AI-training platforms provide the first real workloads. Where their rules permit the planned automation and data handling, those tasks are intended to serve as a paid proving ground for the reusable execution and verification system.
Autonomous execution. Independent verification.
A verification-first task factory
The intended system carries a permitted workload from intake through execution, independent verification, exceptions, and evidence-backed output.
- 01Build direction
External AI-training task
A real software-engineering workload arrives from a permitted source.
- 02Human / policy gate
Policy / permission check
Confirm platform rules, data handling, scope, and authorization before automation.
- 03Build direction
Task intake
Capture instructions, repository context, deliverables, and acceptance requirements.
- 04Build direction
Normalize
Translate the workload into a stable internal task representation.
- 05Build direction
Plan
Form an execution plan and identify verification and review gates.
- 06Build direction
Autonomous execution
Perform supported engineering work without treating unsupported cases as automated.
- 07Demonstrated
Verification
Evaluate the result independently from the solver. Demo V0 proves this layer.
- 08Build direction
Critic / retry
Check failure modes and shortcuts; retry when the evidence supports another attempt.
- 09Human / policy gate
Human review if required
Escalate ambiguity, subjective criteria, compliance decisions, and required approval.
- 10Build direction
Verified output
Package the result, evidence, and unresolved exceptions for delivery.
Demo V0 proves that one candidate repair can be evaluated against a frozen task, discriminated from no repair, and packaged with auditable evidence.
The public record does not prove external task intake, policy checks, autonomous planning or execution, critic / retry behavior, or direct workload delivery.
Why begin with existing AI-training platforms?
They already contain real economic demand, concrete instructions, and acceptance requirements. That makes them a useful place to validate the system against reality.
- 01
Real demand
The workload exists because someone already values the completed task.
- 02
Real acceptance criteria
Instructions and rejection conditions are part of the task, not invented for a demo.
- 03
Real task distributions
Varied repositories and requirements expose where intake and planning break down.
- 04
Meaningful feedback
Accepted and rejected work provides an external signal for improving the system.
- 05
Early revenue
Paid tasks can help fund development without pretending the marketplace is the final product.
- 06
Reusable patterns
Repeated workload families reveal which capabilities should become shared infrastructure.
Marketplace tasks are the input and proving ground. FaultWright is the reusable system being built.
Automation and data handling must remain within the rules and permissions of each platform, customer, and task.
Each completed task should leave reusable capability behind
The goal is not only to finish one workload. Repeated tasks should improve the shared system used for the next supported workload.
- Task intake
- Task classification
- Repository / environment setup
- Execution planning
- Solver orchestration
- Clean replay
- Verification
- Critic / anti-shortcut checks
- Retry logic
- Evidence generation
- Result packaging
Early work may look task-by-task from the outside. The intended structural difference is what becomes reusable after each task.
- Execution logic becomes reusable software.
- Human intervention should decline across repeated, supported task families.
- Independent verification is part of the system.
- Each outcome can improve shared infrastructure.
From paid proving ground to direct task streams
This is the current market-entry horizon, not a claim that all three stages exist today.
- Stage 1Near-term entry
Paid platform tasks
Use permitted real workloads for external feedback and early revenue. No paying customers are claimed today.
- Stage 2Build direction
Reusable autonomous execution
Reduce human intervention across repeated, supported task families while keeping verification independent.
- Stage 3Future target
Direct task streams
Allow customers or platforms to send supported workloads directly into FaultWright. This does not exist today.
Demo V0 — Verification Core
The public demo below focuses on one layer of FaultWright: proving that a candidate software repair can be evaluated against a frozen task and independently verified.
It demonstrates discriminative verification, canonical outcomes, evidence references, artifact integrity, and provenance.
It does not demonstrate autonomous intake or execution of an external platform task. Those capabilities remain the current build direction.
How Demo V0 evaluates a repair
Six steps, all bound to one frozen task identity. This is the verification layer of the broader system, and its output is a behavioral outcome rather than an opinion about the diff.
- 01
Known-good software
Start from software whose expected behavior is already established and observable.
- 02
Controlled fault
Introduce one specific, reproducible failure. The task identity is frozen at this point.
task IF-WEBHOOK-00001
- 03
Verified failure
Confirm the fault is observable before any repair is attempted. An undetectable fault cannot be evaluated.
environment webhook 1.0.0
- 04
Repair attempt
Submit a candidate repair against the same frozen task and the same environment.
2 controlled attempts in this record
- 05
Behavioral verification
Run verification again. The question is whether expected behavior is restored, not whether the diff looks plausible.
visible tests: test/visible.js
- 06
Evaluation result
Record a canonical outcome and a capability verdict together with the evidence that produced them.
PASS · SOLVED / FAIL · NOT_SOLVED
Webhook idempotency: duplicate deliveries must not duplicate invoices
A webhook provider may deliver the same event more than once. Correct software must handle those retries without creating duplicate invoices.
- Starting condition
- Webhook processor accepting payment event deliveries from a provider that may retry them.
- artifact · task.public_summary.starting_state
- Expected behavior
- Multiple deliveries of the same provider event must produce at most one invoice per event. Retries may arrive sequentially or overlap concurrently. Keep the public API unchanged.
- artifact · task.public_summary.problem_statement
- Failure condition
- Repeated deliveries of one event are treated as distinct events, so a single payment can produce more than one invoice.
- editorial · negation of the expected behavior
- Verification target
- Deliver the same event repeatedly, sequentially and concurrently, and confirm that at most one invoice exists per event while the public API is unchanged.
- artifact · visible tests: test/visible.js
Engine narrative — 5 statements exported with the record
- 01This is representative of coding/evaluation work performed in the AI-training market.
- 02FaultFoundry bound that requirement to a reproducible software task.
- 03Two repair attempts were evaluated under the same task/environment.
- 04FaultFoundry's existing evaluation system determined the canonical outcomes.
- 05The resulting evidence was packaged for human review.
Two controls. One frozen task.
These controls verify that the evaluation pipeline distinguishes a valid repair from no repair under the same frozen task and environment. They are deterministic references, not competing AI models.
Reference Repair
Positive control. A repair already known to restore the expected behavior is replayed onto the faulted software.
PASS · SOLVED
A working pipeline must accept it.
PASS · SOLVED
Repair accepted · verification OK
- Submission
- Known-good patch applied
- Profileprofile_id
replay-known-good- Solversolver_id
ref-overlay-replay:known-good- Attemptattempt_id
51a857ff083a- Patchpatch_sha256
94b5e86dcc…715ca99
No Repair
Negative control. Nothing is changed, so the fault stays in place and verification runs against the broken software.
FAIL · NOT_SOLVED
A working pipeline must reject it.
FAIL · NOT_SOLVED
Repair rejected · verification FAILED
- Submission
- No patch submitted
- Profileprofile_id
noop- Solversolver_id
ref-noop- Attemptattempt_id
d31aa52ff859- Patchpatch_sha256
- No patch submitted
The controls validate the evaluator, not the other way round.
Both controls use the same task identity and environment. Only the repair behavior changes. If the known-good repair had not passed, or the no-op attempt had passed, the evaluation pipeline itself would be suspect.
Shared identity: task IF-WEBHOOK-00001 · environment webhook 1.0.0 · fingerprint 13dbadc6af…479c614
- Consistent
Reference Repair
Positive control
expected PASS · SOLVED
observed PASS · SOLVED
- Consistent
No Repair
Negative control
expected FAIL · NOT_SOLVED
observed FAIL · NOT_SOLVED
Pipeline discriminatesThe evaluator accepted the valid repair and rejected the absent one under identical conditions.
Verification record
Every value below is read from the frozen artifact. Digests are shown in full so they can be compared byte for byte with the raw file.
Run
Identity and provenance of this evaluation run.
- Run IDrun_id
879c077fcda2ee530deb9d7f7ad290a3- Createdcreated_at
- 2026-09-28 12:11 UTC (2026-09-28T12:11:07.243873+00:00)
- Schemaschema
demo-v0- Scenarioscenario.scenario_id
coding-ai-evaluation-webhook- Categoryscenario.market_category
- Coding AI evaluation
- Descriptionscenario.short_description
- A reproducible software repair, evaluated under one controlled environment, with the outcome taken from the existing verifier.
Task identity
The frozen task and environment shared by every attempt.
- Task IDtask.task_id
IF-WEBHOOK-00001- Titletask.public_summary.title
- Webhook idempotency: duplicate deliveries must not duplicate invoices
- Environmentenvironment_id · environment_version
webhook 1.0.0- Task fingerprinttask.fingerprint
13dbadc6af39170b067b55adabded22c0ffe58d6417a3830f78a15888479c614- Environment digesttask.environment_digest
a545943ef03516d9e3cded8de79d94b8053305ead8d58129959bf8b85b6eef63- Visible testspublic_summary.visible_tests
test/visible.js
Attempt · Reference Repair
Positive control. A repair already known to restore the expected behavior is replayed onto the faulted software.
- Attempt IDevidence_refs[].attempt_id
51a857ff083a- Profileprofile_id
replay-known-good- Solversolver_id
ref-overlay-replay:known-good- Outcomecanonical_outcome
- PASS
- Verdictcanonical_capability_verdict
- SOLVED
- Verification phaseverification_phase
- OK
- Patch SHA-256patch_sha256
94b5e86dcc4c36e0ae0cbfa47e26c1e790998f3665fc82363d1c8994a715ca99- Summaryhuman_summary
- Repair accepted
Attempt · No Repair
Negative control. Nothing is changed, so the fault stays in place and verification runs against the broken software.
- Attempt IDevidence_refs[].attempt_id
d31aa52ff859- Profileprofile_id
noop- Solversolver_id
ref-noop- Outcomecanonical_outcome
- FAIL
- Verdictcanonical_capability_verdict
- NOT_SOLVED
- Verification phaseverification_phase
- FAILED
- Patch SHA-256patch_sha256
- Empty — no patch was submitted
- Summaryhuman_summary
- Repair rejected
Evidence references
5 references exported with the record, in artifact order.
- 01
Task fingerprint
task_fingerprint
13dbadc6af39170b067b55adabded22c0ffe58d6417a3830f78a15888479c614 - 02
Environment execution digest
environment_execution_digest
a545943ef03516d9e3cded8de79d94b8053305ead8d58129959bf8b85b6eef63 - 03
Attempt record
attempt
- attempt_id
51a857ff083a- profile_id
replay-known-good- solver_id
ref-overlay-replay:known-good
- 04
Patch digest
patch_sha256
94b5e86dcc4c36e0ae0cbfa47e26c1e790998f3665fc82363d1c8994a715ca99 - 05
Attempt record
attempt
- attempt_id
d31aa52ff859- profile_id
noop- solver_id
ref-noop
A frozen record, not a live computation
The displayed result is rendered from a frozen evaluation artifact rather than recomputed in the browser. Both files are published unmodified and can be hashed independently.
demo-v0.json
Task identity, attempts, canonical outcomes, evidence references, and the engine narrative.
- Published at
/demo/demo-v0.json- Size
- 3,319 bytes · LF newlines
- SHA-256
c5671ad60f8a0c4628d463d5daf7ee1d150e885291dc93b8bff7e57c5605a102
demo-v0-freeze.json
Digest and size of the record, the engine commit that exported it, and three provenance timestamps.
- Published at
/demo/demo-v0-freeze.json- Size
- 2,174 bytes · LF newlines
- SHA-256
71fe5d64753a10375207921f3c4dfdbba10f255147f6b8859b43b7b70b7795d7- Artifact createdartifact_created_at
- 2026-09-28 12:11:07 UTC
- Manifest generatedmanifest_generated_at
- 2026-09-28 12:11:59 UTC
- Manifest correctedmanifest_corrected_at
- 2026-09-28 17:31:52 UTC
- Engine commitdemo_implementation_commit
3f8229244954a3308f6d23ebf8c317aa779230f6
The evaluation was not recomputed.
Three timestamps are kept apart on purpose: when the record was exported, when the engine generated the manifest, and when the published digest was corrected after newline normalization. The outcomes, identifiers and evidence date from the first.
correction.reason — Newline normalization. demo_json_sha256, demo_json_bytes, demo_json_newline and demo_json_path were recomputed over the canonical UTF-8, LF-terminated bytes of the published record. The evaluation record (demo-v0.json) was not re-run or edited; its content is as exported at artifact_created_at.
- 01
Evaluation record exported
The engine exported run
879c077fcda2with its outcomes and evidence. This is the record’s owncreated_at.artifact_created_at
- 02
Manifest generated by the engine
Original digest, computed over a CRLF export of the same content:
03d9fedaa8…76fd1b0manifest_generated_at
- 03
Manifest digest corrected
Only manifest fields changed; the record was not re-run or edited. Digest recomputed over the canonical LF bytes:
c5671ad60f…605a102manifest_corrected_at
The build hashes what it ships.
During the production build, the site reads the served copy of demo-v0.json, computes its SHA-256, and compares it with the manifest. Any mismatch fails the build. Newlines are normalized to LF by .gitattributes, so the digest is the same on every operating system and on the deployment host.
- Manifest digestdemo_json_sha256
c5671ad60f8a0c4628d463d5daf7ee1d150e885291dc93b8bff7e57c5605a102- Computed digestsha256(served bytes)
c5671ad60f8a0c4628d463d5daf7ee1d150e885291dc93b8bff7e57c5605a102- Size
- 3,319 bytes · manifest 3,319 bytes · LF
- Result
- Digests match
Two commands.
Download the record and hash it locally. The digest must equal the manifest value below.
macOS / Linux
curl -sO https://fault-wright-demo.vercel.app/demo/demo-v0.json
shasum -a 256 demo-v0.jsonWindows PowerShell
curl.exe -sO https://fault-wright-demo.vercel.app/demo/demo-v0.json
Get-FileHash -Algorithm SHA256 demo-v0.jsonExpected digest
c5671ad60f8a0c4628d463d5daf7ee1d150e885291dc93b8bff7e57c5605a102Preview demo-v0.json — exact bytes as served, 3,319 bytes
{
"attempts": [
{
"canonical_capability_verdict": "SOLVED",
"canonical_outcome": "PASS",
"human_summary": "Repair accepted",
"patch_sha256": "94b5e86dcc4c36e0ae0cbfa47e26c1e790998f3665fc82363d1c8994a715ca99",
"profile_id": "replay-known-good",
"solver_id": "ref-overlay-replay:known-good",
"solver_label": "Reference Repair",
"verification_phase": "OK"
},
{
"canonical_capability_verdict": "NOT_SOLVED",
"canonical_outcome": "FAIL",
"human_summary": "Repair rejected",
"patch_sha256": "",
"profile_id": "noop",
"solver_id": "ref-noop",
"solver_label": "No Repair",
"verification_phase": "FAILED"
}
],
"created_at": "2026-09-28T12:11:07.243873+00:00",
"evidence_refs": [
{
"kind": "task_fingerprint",
"sha256": "13dbadc6af39170b067b55adabded22c0ffe58d6417a3830f78a15888479c614"
},
{
"kind": "environment_execution_digest",
"sha256": "a545943ef03516d9e3cded8de79d94b8053305ead8d58129959bf8b85b6eef63"
},
{
"attempt_id": "51a857ff083a",
"kind": "attempt",
"profile_id": "replay-known-good",
"solver_id": "ref-overlay-replay:known-good"
},
{
"kind": "patch_sha256",
"sha256": "94b5e86dcc4c36e0ae0cbfa47e26c1e790998f3665fc82363d1c8994a715ca99"
},
{
"attempt_id": "d31aa52ff859",
"kind": "attempt",
"profile_id": "noop",
"solver_id": "ref-noop"
}
],
"narrative": [
"This is representative of coding/evaluation work performed in the AI-training market.",
"FaultFoundry bound that requirement to a reproducible software task.",
"Two repair attempts were evaluated under the same task/environment.",
"FaultFoundry's existing evaluation system determined the canonical outcomes.",
"The resulting evidence was packaged for human review."
],
"run_id": "879c077fcda2ee530deb9d7f7ad290a3",
"scenario": {
"expected_deliverables": [
"reproducible_problem",
"controlled_environment",
"verified_solution_attempt",
"discriminative_verification",
"evidence_package"
],
"market_category": "coding_ai_evaluation",
"scenario_id": "coding-ai-evaluation-webhook",
"short_description": "A reproducible software repair, evaluated under one controlled environment, with the outcome taken from the existing verifier."
},
"schema": "demo-v0",
"task": {
"environment_digest": "a545943ef03516d9e3cded8de79d94b8053305ead8d58129959bf8b85b6eef63",
"environment_id": "webhook",
"environment_version": "1.0.0",
"fingerprint": "13dbadc6af39170b067b55adabded22c0ffe58d6417a3830f78a15888479c614",
"public_summary": {
"environment_id": "webhook",
"problem_statement": "Multiple deliveries of the same provider event must produce at most one invoice per event. Retries may arrive sequentially or overlap concurrently. Keep the public API unchanged.",
"starting_state": "Webhook processor accepting payment event deliveries from a provider that may retry them.",
"task_id": "IF-WEBHOOK-00001",
"title": "Webhook idempotency: duplicate deliveries must not duplicate invoices",
"visible_tests": [
"test/visible.js"
]
},
"task_id": "IF-WEBHOOK-00001"
}
}
What this record does and does not claim
A controlled proof comes before broader claims. The deliverables on the left are the ones the artifact itself lists.
Demonstrated
scenario.expected_deliverables
- 01
Reproducible problem
A failure that can be reproduced on demand from a frozen task identity.
- 02
Controlled environment
One pinned environment, identified by version and execution digest.
- 03
Verified solution attempt
A repair attempt whose effect is checked by re-running verification, not by reading the diff.
- 04
Discriminative verification
A verifier shown to accept a valid repair and reject no repair under identical conditions.
- 05
Evidence package
Identifiers, digests, and outcomes exported as a public, hashable artifact.
Not claimed
editorial
This record does not prove autonomous intake, planning, or execution of an external platform task.
It does not imply that every platform or software-engineering task family is supported.
No AI model or coding agent was evaluated in this record. Both attempts are deterministic controls.
The controls are not competing systems. They exist to check the evaluator, not to rank solvers.
No success rates, benchmark scores, customers, or production usage are asserted.
- Demo V0 proves the verification core under one frozen task and environment.
- Autonomous external task intake and execution are under active development.
- Not every platform or software-engineering task family is currently supported.
- Automation and data handling must be permitted by the platform, customer, and task.
- Humans remain necessary for ambiguity, subjective criteria, compliance decisions, exceptions, and required approval.