Skip to content
FaultWright
FaultWright · Current build direction

Real tasks in. Verified results out.

FaultWright is being built to autonomously execute permitted software-engineering AI-training tasks and independently verify the result.

Existing AI-training platforms provide the first real workloads. Where their rules permit the planned automation and data handling, those tasks are intended to serve as a paid proving ground for the reusable execution and verification system.

Autonomous execution. Independent verification.

Current build direction

A verification-first task factory

The intended system carries a permitted workload from intake through execution, independent verification, exceptions, and evidence-backed output.

Current build direction · end-to-end flowCurrently demonstrated · verification core
  1. 01Build direction

    External AI-training task

    A real software-engineering workload arrives from a permitted source.

  2. 02Human / policy gate

    Policy / permission check

    Confirm platform rules, data handling, scope, and authorization before automation.

  3. 03Build direction

    Task intake

    Capture instructions, repository context, deliverables, and acceptance requirements.

  4. 04Build direction

    Normalize

    Translate the workload into a stable internal task representation.

  5. 05Build direction

    Plan

    Form an execution plan and identify verification and review gates.

  6. 06Build direction

    Autonomous execution

    Perform supported engineering work without treating unsupported cases as automated.

  7. 07Demonstrated

    Verification

    Evaluate the result independently from the solver. Demo V0 proves this layer.

  8. 08Build direction

    Critic / retry

    Check failure modes and shortcuts; retry when the evidence supports another attempt.

  9. 09Human / policy gate

    Human review if required

    Escalate ambiguity, subjective criteria, compliance decisions, and required approval.

  10. 10Build direction

    Verified output

    Package the result, evidence, and unresolved exceptions for delivery.

Currently demonstrated

Demo V0 proves that one candidate repair can be evaluated against a frozen task, discriminated from no repair, and packaged with auditable evidence.

Not proved by Demo V0

The public record does not prove external task intake, policy checks, autonomous planning or execution, critic / retry behavior, or direct workload delivery.

Paid proving ground

Why begin with existing AI-training platforms?

They already contain real economic demand, concrete instructions, and acceptance requirements. That makes them a useful place to validate the system against reality.

  1. 01

    Real demand

    The workload exists because someone already values the completed task.

  2. 02

    Real acceptance criteria

    Instructions and rejection conditions are part of the task, not invented for a demo.

  3. 03

    Real task distributions

    Varied repositories and requirements expose where intake and planning break down.

  4. 04

    Meaningful feedback

    Accepted and rejected work provides an external signal for improving the system.

  5. 05

    Early revenue

    Paid tasks can help fund development without pretending the marketplace is the final product.

  6. 06

    Reusable patterns

    Repeated workload families reveal which capabilities should become shared infrastructure.

Marketplace tasks are the input and proving ground. FaultWright is the reusable system being built.

Automation and data handling must remain within the rules and permissions of each platform, customer, and task.

What accumulates

Each completed task should leave reusable capability behind

The goal is not only to finish one workload. Repeated tasks should improve the shared system used for the next supported workload.

  • Task intake
  • Task classification
  • Repository / environment setup
  • Execution planning
  • Solver orchestration
  • Clean replay
  • Verification
  • Critic / anti-shortcut checks
  • Retry logic
  • Evidence generation
  • Result packaging
Service vs. product direction

Early work may look task-by-task from the outside. The intended structural difference is what becomes reusable after each task.

Traditional task execution
  • Similar work is performed again for each task.
  • Capacity grows mainly by adding human time.
  • Process knowledge may remain informal.
FaultWright direction
  • Execution logic becomes reusable software.
  • Human intervention should decline across repeated, supported task families.
  • Independent verification is part of the system.
  • Each outcome can improve shared infrastructure.
Near-term roadmap

From paid proving ground to direct task streams

This is the current market-entry horizon, not a claim that all three stages exist today.

  1. Stage 1Near-term entry

    Paid platform tasks

    Use permitted real workloads for external feedback and early revenue. No paying customers are claimed today.

  2. Stage 2Build direction

    Reusable autonomous execution

    Reduce human intervention across repeated, supported task families while keeping verification independent.

  3. Stage 3Future target

    Direct task streams

    Allow customers or platforms to send supported workloads directly into FaultWright. This does not exist today.

Currently demonstrated

Demo V0 — Verification Core

The public demo below focuses on one layer of FaultWright: proving that a candidate software repair can be evaluated against a frozen task and independently verified.

It demonstrates discriminative verification, canonical outcomes, evidence references, artifact integrity, and provenance.

It does not demonstrate autonomous intake or execution of an external platform task. Those capabilities remain the current build direction.

01Verification core method

How Demo V0 evaluates a repair

Six steps, all bound to one frozen task identity. This is the verification layer of the broader system, and its output is a behavioral outcome rather than an opinion about the diff.

  1. 01

    Known-good software

    Start from software whose expected behavior is already established and observable.

  2. 02

    Controlled fault

    Introduce one specific, reproducible failure. The task identity is frozen at this point.

    task IF-WEBHOOK-00001

  3. 03

    Verified failure

    Confirm the fault is observable before any repair is attempted. An undetectable fault cannot be evaluated.

    environment webhook 1.0.0

  4. 04

    Repair attempt

    Submit a candidate repair against the same frozen task and the same environment.

    2 controlled attempts in this record

  5. 05

    Behavioral verification

    Run verification again. The question is whether expected behavior is restored, not whether the diff looks plausible.

    visible tests: test/visible.js

  6. 06

    Evaluation result

    Record a canonical outcome and a capability verdict together with the evidence that produced them.

    PASS · SOLVED / FAIL · NOT_SOLVED

02Frozen challenge · Webhook idempotency

Webhook idempotency: duplicate deliveries must not duplicate invoices

A webhook provider may deliver the same event more than once. Correct software must handle those retries without creating duplicate invoices.

Starting condition
Webhook processor accepting payment event deliveries from a provider that may retry them.
artifact · task.public_summary.starting_state
Expected behavior
Multiple deliveries of the same provider event must produce at most one invoice per event. Retries may arrive sequentially or overlap concurrently. Keep the public API unchanged.
artifact · task.public_summary.problem_statement
Failure condition
Repeated deliveries of one event are treated as distinct events, so a single payment can produce more than one invoice.
editorial · negation of the expected behavior
Verification target
Deliver the same event repeatedly, sequentially and concurrently, and confirm that at most one invoice exists per event while the public API is unchanged.
artifact · visible tests: test/visible.js
Engine narrative — 5 statements exported with the record
  1. 01This is representative of coding/evaluation work performed in the AI-training market.
  2. 02FaultFoundry bound that requirement to a reproducible software task.
  3. 03Two repair attempts were evaluated under the same task/environment.
  4. 04FaultFoundry's existing evaluation system determined the canonical outcomes.
  5. 05The resulting evidence was packaged for human review.
03Pipeline controls

Two controls. One frozen task.

These controls verify that the evaluation pipeline distinguishes a valid repair from no repair under the same frozen task and environment. They are deterministic references, not competing AI models.

Control 01 / 02 · positive control

Reference Repair

Positive control. A repair already known to restore the expected behavior is replayed onto the faulted software.

Accepted
Expected

PASS · SOLVED

A working pipeline must accept it.

Observed

PASS · SOLVED

Repair accepted · verification OK

Submission
Known-good patch applied
Profileprofile_id
replay-known-good
Solversolver_id
ref-overlay-replay:known-good
Attemptattempt_id
51a857ff083a
Patchpatch_sha256
94b5e86dcc…715ca99
Control 02 / 02 · negative control

No Repair

Negative control. Nothing is changed, so the fault stays in place and verification runs against the broken software.

Rejected
Expected

FAIL · NOT_SOLVED

A working pipeline must reject it.

Observed

FAIL · NOT_SOLVED

Repair rejected · verification FAILED

Submission
No patch submitted
Profileprofile_id
noop
Solversolver_id
ref-noop
Attemptattempt_id
d31aa52ff859
Patchpatch_sha256
No patch submitted
Why two controls

The controls validate the evaluator, not the other way round.

Both controls use the same task identity and environment. Only the repair behavior changes. If the known-good repair had not passed, or the no-op attempt had passed, the evaluation pipeline itself would be suspect.

Shared identity: task IF-WEBHOOK-00001 · environment webhook 1.0.0 · fingerprint 13dbadc6af…479c614

  • Reference Repair

    Positive control

    Consistent

    expected PASS · SOLVED

    observed PASS · SOLVED

  • No Repair

    Negative control

    Consistent

    expected FAIL · NOT_SOLVED

    observed FAIL · NOT_SOLVED

Pipeline discriminatesThe evaluator accepted the valid repair and rejected the absent one under identical conditions.

04Technical evidence

Verification record

Every value below is read from the frozen artifact. Digests are shown in full so they can be compared byte for byte with the raw file.

Run

Identity and provenance of this evaluation run.

Run IDrun_id
879c077fcda2ee530deb9d7f7ad290a3
Createdcreated_at
2026-09-28 12:11 UTC (2026-09-28T12:11:07.243873+00:00)
Schemaschema
demo-v0
Scenarioscenario.scenario_id
coding-ai-evaluation-webhook
Categoryscenario.market_category
Coding AI evaluation
Descriptionscenario.short_description
A reproducible software repair, evaluated under one controlled environment, with the outcome taken from the existing verifier.

Task identity

The frozen task and environment shared by every attempt.

Task IDtask.task_id
IF-WEBHOOK-00001
Titletask.public_summary.title
Webhook idempotency: duplicate deliveries must not duplicate invoices
Environmentenvironment_id · environment_version
webhook 1.0.0
Task fingerprinttask.fingerprint
13dbadc6af39170b067b55adabded22c0ffe58d6417a3830f78a15888479c614
Environment digesttask.environment_digest
a545943ef03516d9e3cded8de79d94b8053305ead8d58129959bf8b85b6eef63
Visible testspublic_summary.visible_tests
  • test/visible.js

Attempt · Reference Repair

Positive control. A repair already known to restore the expected behavior is replayed onto the faulted software.

Attempt IDevidence_refs[].attempt_id
51a857ff083a
Profileprofile_id
replay-known-good
Solversolver_id
ref-overlay-replay:known-good
Outcomecanonical_outcome
PASS
Verdictcanonical_capability_verdict
SOLVED
Verification phaseverification_phase
OK
Patch SHA-256patch_sha256
94b5e86dcc4c36e0ae0cbfa47e26c1e790998f3665fc82363d1c8994a715ca99
Summaryhuman_summary
Repair accepted

Attempt · No Repair

Negative control. Nothing is changed, so the fault stays in place and verification runs against the broken software.

Attempt IDevidence_refs[].attempt_id
d31aa52ff859
Profileprofile_id
noop
Solversolver_id
ref-noop
Outcomecanonical_outcome
FAIL
Verdictcanonical_capability_verdict
NOT_SOLVED
Verification phaseverification_phase
FAILED
Patch SHA-256patch_sha256
Empty — no patch was submitted
Summaryhuman_summary
Repair rejected

Evidence references

5 references exported with the record, in artifact order.

  1. 01

    Task fingerprint

    task_fingerprint

    13dbadc6af39170b067b55adabded22c0ffe58d6417a3830f78a15888479c614
  2. 02

    Environment execution digest

    environment_execution_digest

    a545943ef03516d9e3cded8de79d94b8053305ead8d58129959bf8b85b6eef63
  3. 03

    Attempt record

    attempt

    attempt_id
    51a857ff083a
    profile_id
    replay-known-good
    solver_id
    ref-overlay-replay:known-good
  4. 04

    Patch digest

    patch_sha256

    94b5e86dcc4c36e0ae0cbfa47e26c1e790998f3665fc82363d1c8994a715ca99
  5. 05

    Attempt record

    attempt

    attempt_id
    d31aa52ff859
    profile_id
    noop
    solver_id
    ref-noop
05Artifacts

A frozen record, not a live computation

The displayed result is rendered from a frozen evaluation artifact rather than recomputed in the browser. Both files are published unmodified and can be hashed independently.

Evaluation record

demo-v0.json

Task identity, attempts, canonical outcomes, evidence references, and the engine narrative.

Verified at build
Published at
/demo/demo-v0.json
Size
3,319 bytes · LF newlines
SHA-256
c5671ad60f8a0c4628d463d5daf7ee1d150e885291dc93b8bff7e57c5605a102
Open raw fileDownload/demo/demo-v0.json
Freeze manifest

demo-v0-freeze.json

Digest and size of the record, the engine commit that exported it, and three provenance timestamps.

Published at
/demo/demo-v0-freeze.json
Size
2,174 bytes · LF newlines
SHA-256
71fe5d64753a10375207921f3c4dfdbba10f255147f6b8859b43b7b70b7795d7
Artifact createdartifact_created_at
2026-09-28 12:11:07 UTC
Manifest generatedmanifest_generated_at
2026-09-28 12:11:59 UTC
Manifest correctedmanifest_corrected_at
2026-09-28 17:31:52 UTC
Engine commitdemo_implementation_commit
3f8229244954a3308f6d23ebf8c317aa779230f6
Open raw fileDownload/demo/demo-v0-freeze.json
Provenance

The evaluation was not recomputed.

Three timestamps are kept apart on purpose: when the record was exported, when the engine generated the manifest, and when the published digest was corrected after newline normalization. The outcomes, identifiers and evidence date from the first.

correction.reason — Newline normalization. demo_json_sha256, demo_json_bytes, demo_json_newline and demo_json_path were recomputed over the canonical UTF-8, LF-terminated bytes of the published record. The evaluation record (demo-v0.json) was not re-run or edited; its content is as exported at artifact_created_at.

  1. 01

    Evaluation record exported

    The engine exported run 879c077fcda2 with its outcomes and evidence. This is the record’s own created_at.

    artifact_created_at

  2. 02

    Manifest generated by the engine

    Original digest, computed over a CRLF export of the same content: 03d9fedaa8…76fd1b0

    manifest_generated_at

  3. 03

    Manifest digest corrected

    Only manifest fields changed; the record was not re-run or edited. Digest recomputed over the canonical LF bytes: c5671ad60f…605a102

    manifest_corrected_at

Build-time verification

The build hashes what it ships.

During the production build, the site reads the served copy of demo-v0.json, computes its SHA-256, and compares it with the manifest. Any mismatch fails the build. Newlines are normalized to LF by .gitattributes, so the digest is the same on every operating system and on the deployment host.

Manifest digestdemo_json_sha256
c5671ad60f8a0c4628d463d5daf7ee1d150e885291dc93b8bff7e57c5605a102
Computed digestsha256(served bytes)
c5671ad60f8a0c4628d463d5daf7ee1d150e885291dc93b8bff7e57c5605a102
Size
3,319 bytes · manifest 3,319 bytes · LF
Result
Digests match
Verify it yourself

Two commands.

Download the record and hash it locally. The digest must equal the manifest value below.

macOS / Linux

curl -sO https://fault-wright-demo.vercel.app/demo/demo-v0.json
shasum -a 256 demo-v0.json

Windows PowerShell

curl.exe -sO https://fault-wright-demo.vercel.app/demo/demo-v0.json
Get-FileHash -Algorithm SHA256 demo-v0.json

Expected digest

c5671ad60f8a0c4628d463d5daf7ee1d150e885291dc93b8bff7e57c5605a102
Preview demo-v0.json — exact bytes as served, 3,319 bytes
{
  "attempts": [
    {
      "canonical_capability_verdict": "SOLVED",
      "canonical_outcome": "PASS",
      "human_summary": "Repair accepted",
      "patch_sha256": "94b5e86dcc4c36e0ae0cbfa47e26c1e790998f3665fc82363d1c8994a715ca99",
      "profile_id": "replay-known-good",
      "solver_id": "ref-overlay-replay:known-good",
      "solver_label": "Reference Repair",
      "verification_phase": "OK"
    },
    {
      "canonical_capability_verdict": "NOT_SOLVED",
      "canonical_outcome": "FAIL",
      "human_summary": "Repair rejected",
      "patch_sha256": "",
      "profile_id": "noop",
      "solver_id": "ref-noop",
      "solver_label": "No Repair",
      "verification_phase": "FAILED"
    }
  ],
  "created_at": "2026-09-28T12:11:07.243873+00:00",
  "evidence_refs": [
    {
      "kind": "task_fingerprint",
      "sha256": "13dbadc6af39170b067b55adabded22c0ffe58d6417a3830f78a15888479c614"
    },
    {
      "kind": "environment_execution_digest",
      "sha256": "a545943ef03516d9e3cded8de79d94b8053305ead8d58129959bf8b85b6eef63"
    },
    {
      "attempt_id": "51a857ff083a",
      "kind": "attempt",
      "profile_id": "replay-known-good",
      "solver_id": "ref-overlay-replay:known-good"
    },
    {
      "kind": "patch_sha256",
      "sha256": "94b5e86dcc4c36e0ae0cbfa47e26c1e790998f3665fc82363d1c8994a715ca99"
    },
    {
      "attempt_id": "d31aa52ff859",
      "kind": "attempt",
      "profile_id": "noop",
      "solver_id": "ref-noop"
    }
  ],
  "narrative": [
    "This is representative of coding/evaluation work performed in the AI-training market.",
    "FaultFoundry bound that requirement to a reproducible software task.",
    "Two repair attempts were evaluated under the same task/environment.",
    "FaultFoundry's existing evaluation system determined the canonical outcomes.",
    "The resulting evidence was packaged for human review."
  ],
  "run_id": "879c077fcda2ee530deb9d7f7ad290a3",
  "scenario": {
    "expected_deliverables": [
      "reproducible_problem",
      "controlled_environment",
      "verified_solution_attempt",
      "discriminative_verification",
      "evidence_package"
    ],
    "market_category": "coding_ai_evaluation",
    "scenario_id": "coding-ai-evaluation-webhook",
    "short_description": "A reproducible software repair, evaluated under one controlled environment, with the outcome taken from the existing verifier."
  },
  "schema": "demo-v0",
  "task": {
    "environment_digest": "a545943ef03516d9e3cded8de79d94b8053305ead8d58129959bf8b85b6eef63",
    "environment_id": "webhook",
    "environment_version": "1.0.0",
    "fingerprint": "13dbadc6af39170b067b55adabded22c0ffe58d6417a3830f78a15888479c614",
    "public_summary": {
      "environment_id": "webhook",
      "problem_statement": "Multiple deliveries of the same provider event must produce at most one invoice per event. Retries may arrive sequentially or overlap concurrently. Keep the public API unchanged.",
      "starting_state": "Webhook processor accepting payment event deliveries from a provider that may retry them.",
      "task_id": "IF-WEBHOOK-00001",
      "title": "Webhook idempotency: duplicate deliveries must not duplicate invoices",
      "visible_tests": [
        "test/visible.js"
      ]
    },
    "task_id": "IF-WEBHOOK-00001"
  }
}
06Scope

What this record does and does not claim

A controlled proof comes before broader claims. The deliverables on the left are the ones the artifact itself lists.

Demonstrated

scenario.expected_deliverables

  1. 01

    Reproducible problem

    A failure that can be reproduced on demand from a frozen task identity.

  2. 02

    Controlled environment

    One pinned environment, identified by version and execution digest.

  3. 03

    Verified solution attempt

    A repair attempt whose effect is checked by re-running verification, not by reading the diff.

  4. 04

    Discriminative verification

    A verifier shown to accept a valid repair and reject no repair under identical conditions.

  5. 05

    Evidence package

    Identifiers, digests, and outcomes exported as a public, hashable artifact.

Not claimed

editorial

  • This record does not prove autonomous intake, planning, or execution of an external platform task.

  • It does not imply that every platform or software-engineering task family is supported.

  • No AI model or coding agent was evaluated in this record. Both attempts are deterministic controls.

  • The controls are not competing systems. They exist to check the evaluator, not to rank solvers.

  • No success rates, benchmark scores, customers, or production usage are asserted.

Honest current state
  • Demo V0 proves the verification core under one frozen task and environment.
  • Autonomous external task intake and execution are under active development.
  • Not every platform or software-engineering task family is currently supported.
  • Automation and data handling must be permitted by the platform, customer, and task.
  • Humans remain necessary for ambiguity, subjective criteria, compliance decisions, exceptions, and required approval.