AI engineering experiment · 8 August 2026

Give AI agents an engineering specification. Measure what they can actually deliver.

EmberScope is an AI engineering experiment: measure how far agents can take a drone-mounted thermal-instrument specification, preserve their calculations and failures, and identify every material point of human direction. The agents perform science and research; people remain responsible for the outputs. Reproducible AI work is evidence for review, not accepted engineering.

Experimental record

The repository records what agents produced and where people changed or bounded the work.

Autonomous agent work

Codex and Claude agents implemented the human-ordered queue through models, optimizer runs, code, tests, CAD, renders, and review pages. Passing a repository test establishes a declared software or artifact check, not engineering acceptance.

Greg Baker’s decisions

Greg set the exact comparison target to 100 mm EFL at F/1.8 on 24 July, directed on 31 July that downloadable CAD was the primary public value, and decides which results are published.

AAO and RFS inputs

The AAO feasibility-study charter sets the research scope. Current mission values are Greg-recorded advisory defaults pending any superseding dated Dani or RFS answer. No RFS operational acceptance is recorded.

Negative results

Five bounded three- through seven-mirror families were run at the common target; zero passed. Their acceptance packets and exact shortfalls remain visible instead of being rewritten as a successful shortlist.

Human responsibility

Greg takes responsibility for publication; AAO or another authorised reviewer must decide technical acceptance, scope, funding, manufacture, and use. No agent can make those decisions.

Roles and responsibility

Agent labour is disclosed; authority and publication responsibility remain human.

Participant Recorded role Authority boundary
OpenAI Codex Autonomous agent work in daily implementation sessions and interactive repository work: models, traces, software, tests, and site artifacts. May execute the ordered research queue; cannot supply AAO/RFS authority or accept an output for a human reviewer.
Anthropic Claude / Claude Code Interactive development and review, including the session that recorded the RFS-011 study defaults. Provides research and implementation assistance, not stakeholder sign-off.
OpenAI GPT / ChatGPT and Google Gemini Background-research drafting and review recorded in the scan-mirror heritage work. Draft or review output is not accepted as evidence without a checkable source or repository artifact.
Greg Baker Sets the charter, records or approves requirement interpretations and course corrections, decides what is published, and reviews the package. Responsible for the published outputs; cannot stand in for an AAO or RFS decision that the project identifies as external.
AAO and RFS inputs The AAO charter defines the feasibility study. A dated Dani or RFS answer can supersede the recorded mission defaults. No RFS operational acceptance, hardware authority, or field-use approval is present in the repository.

The family names describe tools used across the project. Git does not identify the model responsible for every line, so model identity is not used as proof of a technical claim.

Current evidence chain

Each published result can be followed back to versioned source and a reproducible check.

1 · Git history

Commits preserve the source change, author, timestamp, and diff. Dated research notes retain the reported checkpoint instead of silently revising it.

2 · Build revision

scripts/build_static_site.sh writes the Git revision and UTC build time to dist/build.json. Deployment verifies that the remote stamp contains the pushed revision.

3 · Source-spec hash

The canonical prescription JSON is hashed with SHA-256. Layout, CAD, Blender, validation, and review-package exports carry the same source_hash_sha256 value.

4 · Regression gates

Renderer and model tests compare committed artifacts with recalculated outputs. Pushes run Ruff, mypy, ShellCheck, and the full pytest suite; the deployment build separately reruns the site renderers.

5 · Diary and RSS

The dated work diary and its RSS feeds record automation and review sessions. They provide chronology, not a substitute for calculated evidence.

Negative and preliminary results

A failed trace remains part of the evidence.

Daily work may report an unfinished attempt or no committable result. Preliminary engineering is allowed when assumptions and open gates are stated. Process material cannot replace a calculation, trace, test, or human decision that is still absent.

The current P2 shortlist contains five bounded three- through seven-mirror families at 100 mm EFL and F/1.8, and zero passed. The most recent three-mirror result measured 23.5% minimum unobscured throughput, 2294.3 µm worst detector-plane RMS, and 4919.5 µm flat-focus spread. Those figures reject that recorded run against the common gates; they are not a claim that every possible reflective architecture fails.

Independent verification

Clone, test, rebuild, and compare before accepting a claim.

Check Command or comparison What it establishes
Repository tests uv run pytest telescope_explorer/tests The committed code and artifact consistency checks pass in the reviewer’s environment.
Static rebuild ./scripts/build_static_site.sh, then git diff --exit-code -- site The renderer chain reproduces the tracked stakeholder-facing artifacts.
Design export uv run projects/rural-fire-service-camera/scripts/rfs_design_source_of_truth.py The canonical design JSON regenerates the layout, CAD, Blender, validation, and manifest bundle.
Hash comparison Compare source_hash_sha256 across the generated validation and manifest files. The reviewed exports trace to the same canonical prescription.
Published revision Compare the live build.json revision with the Git commit under review. The deployed tree identifies the source revision that produced it.

A reproducible artifact preserves provenance; it does not promote a labelled-preliminary result to a validated instrument. Start from the public repository or return to the project review index for the surrounding evidence.