AAO feasibility study · AI engineering experiment

How far can AI agents take an instrument specification?

EmberScope is the test case: give AI agents an engineering specification for a drone-mounted long-wave infrared fire-spotting payload, measure how far they can take it using committed models and public data, and expose every material point where human direction changed the work. Agents produced calculations, traces, code, CAD, and review files; humans remain responsible for what is published. This is research evidence, not accepted engineering.

What the experiment records

Work, direction, evidence, and acceptance are kept separate.

Autonomous agent work

Codex and Claude agents implemented the ordered queue: models, bounded optimizer runs, source-hashed exports, tests, renders, and public review artifacts. This describes who performed the work, not who has authority to accept it.

Greg Baker’s decisions

Greg set the comparison target to 100 mm EFL at F/1.8 on 24 July and later directed that downloadable CAD become the primary public value. He sets the charter, course corrections, and publication boundary.

AAO and RFS inputs

The AAO feasibility-study charter defines the research task. Mission values are recorded study defaults until a dated Dani or RFS answer supersedes them. No RFS operational acceptance is recorded.

Negative results

Five bounded three- through seven-mirror families were measured at the common target; zero passed the P2 gates. The failed traces, exact shortfalls, and reproducible artifacts remain public.

Human responsibility

People decide whether evidence is credible, whether requirements should change, and whether anything proceeds to manufacture or field use. Publication does not make an AI output accepted engineering.

Project status: design study only — nothing has been built.

EmberScope exists as software and review artifacts: source-hashed CAD review packages, a parametric payload assembly, bounded ray traces, radiometry and pipeline models, and validation plans. No mirrors have been fabricated, no detector has been purchased, and no flights or field measurements have taken place. The current five-candidate optical shortlist has zero passing designs; the CAD is for geometry and package review, not CNC release or fabrication.

Mission problem

The camera has to detect small heat sources without breaking the drone mission.

The working mission is a 150 mm, sub-2 kg thermal payload that can support one-hour revisit planning, controlled false alarms, and replayable evidence for a smouldering target over rural backgrounds.

That makes this a camera-level decision: optics, detector, calibration, data pipeline, payload structure, and field testing have to close together.

Decision focus

AAO should decide whether this is the right optical maturation path.

The current evidence rejects the five bounded optical families at the stated gates. It supports either another declared candidate family or a human decision to stop or reframe the custom-optics path; it does not support fabrication.

Rendered optimized four-mirror EmberScope optical train with traced beam paths inside the 150 millimetre envelope wireframe.
A source-traced four-mirror study artifact from 10 June 2026: gold conic mirrors, three displayed field bundles, the entrance stop ring, and the 150 mm envelope. Later hard-gated work superseded its early validation status; it is not a passing P2 prescription.

Current evidence

The payload model closes its envelope. The optical shortlist does not.

Measured failure

Optical shortlist

The five bounded 100 mm F/1.8 candidates all fail one or more common P2 gates. The latest three-mirror result measures 23.5% minimum throughput, 2294.3 µm worst detector-plane RMS, and 4919.5 µm flat-focus spread.

Modelled

Payload envelope

The parametric assembly measures 148 mm outer span and 1.643922 kg modelled mass, with no axis-aligned component or reserved-clearance violations. It remains an engineering-development model, not a manufacturing release.

Needs proof

Detection performance

GSD, NETD, calibration drift, and false-positive scenes still need measured closure. No hardware exists yet.

Main technical risks

The unresolved issues are specific enough to test.

Optical quality and vignetting

Four ray-path rules are first-class evaluation gates: non-vignetted propagation, small fold angles where practical, detector-plane placement, and flatter image surfaces.

Radiometry and data integrity

Detection claims need raw or pre-automatic-gain-control thermal frames, NUC state, references, and replayable alert records.

Field validation

Bench and surrogate targets must precede controlled burns, and no-fire rural scenes must be measured beside positive detections.

Technical archive

Detailed pages remain available under four grouped paths.

Next AAO decisions

The next review should choose the evidence gates, not just admire the concept render.

1

Mission closure

Confirm target area, altitude authorisation, payload count, false-alarm tolerance, and whether the 150 mm cube is a hard requirement.

2

Optical architecture

Choose the next declared candidate family or stop or reframe the custom-optics path. None of the five failed P2 designs is eligible for tolerancing or vendor review.

3

Detector and radiometry

Lock the detector interface enough to replace generic image-energy assumptions with measured FPA and calibration behaviour.

4

Validation ladder

Approve the bench-to-field route before controlled burn claims or operational readiness claims are made.