Validation plan - 16 May 2026

Dataset validation plan

This is the public plan for eventually proving fire detection without hiding behind random frame splits, visible-flame shortcuts, or unmeasured false positives. The first public radiometric replay now exists, and it rejects the current detector profile on no-fire clutter. The 31 July scope correction retains that result as camera-software evidence, not optical-design evidence.

Evidence base

The plan is grounded in the public FLAME sequence and the current EmberScope mission requirements.

FLAME Pile-burn RGB and thermal-palette data with classification, segmentation, and a documented false-positive imbalance warning.
FLAME 2 Side-by-side RGB/IR prescribed-fire frames with expert fire/no-fire and smoke labels plus burn context records.
FLAME 3 Auxiliary camera-software dataset; its native DJI imagery does not validate an EmberScope optical prescription.
Stress tests Dataset choice is tied to small-target radiometry, false-positive rural scenes, and survey operations.

Implementation checkpoint - 29 July 2026

The FLAME 3 ingestion adapter now rejects incomplete or ambiguous radiometric inputs before scoring.

rfs_flame3_ingest.py scans a caller-supplied dataset directory outside git and implements the published same-stem four-file pairing for raw RGB, corrected-FOV RGB, thermal JPEG, and Celsius TIFF imagery.

Pairing integrity

Every label folder must contain the same filename stems in all four modalities; missing or duplicate stems stop ingestion.

Radiometric gate

Each TIFF is checked as a finite, single-band float32 Celsius raster at the metadata-declared dimensions, with the Celsius-to-Kelvin conversion recorded explicitly.

Provenance

Every source file is SHA-256 checked, and preserved thermal-JPEG EXIF supplies the acquisition day and observed camera model.

Replay contract

The adapter emits a detailed provenance manifest and an RFS-047 dataset-export manifest with group-disjoint burn, site, day, and sensor splits.

The first public FLAME 3 replay is now recorded below. It uses the corrected Thermal/Raw JPG source path and all 738 public Sycan Marsh quartets. It remains a single-burn detector-software diagnostic; the 31 July scope correction closes active RFS-069 work without claiming burn-held-out acceptance.

Source convention: FLAME 3 dataset DOI and the published processing pipeline.

Measured replay - 30 July 2026

The current detector finds the fire-labelled frames but fails the public no-fire stress case.

The unchanged compact-hot-v1 profile was replayed across every radiometric TIFF in the public FLAME 3 Computer Vision Subset (Sycan Marsh), DOI 10.21227/w0mz-aq48. This is real public imagery rather than synthetic injection.

738 Real float32 Celsius TIFFs: 622 Fire and 116 No Fire.
99.4% Frame recall: 618 true positives and 4 false negatives.
6.9% No-fire specificity: 108 false positives and only 8 true negatives.
77,302 Alert packets emitted across the 26 reconstructed sequences.
Frame confusion counts for the single-burn FLAME 3 compact-hot-v1 replay: 618 true positives, 108 false positives, 8 true negatives, and 4 false negatives.

All four reconstructed no-fire sequences produced alerts. The high recall is therefore not an operational success: broad fire-scene labels and class imbalance conceal a clutter-rejection failure that would overload review. The four missed fire frames are one-frame sequences and cannot satisfy the configured two-frame persistence gate.

The reproducible result manifest records the exact archive, code and manifest checksums; unchanged thresholds; declared split and sequence rule; frame and sequence metrics; every failed frame ID; and the source limitations.

This is a single-burn diagnostic, not burn-held-out acceptance evidence. The public source contains one Sycan Marsh burn; the paper says the full six-burn source is available on request. Preserved thermal-JPEG EXIF also reports 25-27 October 2022 while the public data card describes 25-27 October 2023; the replay preserves the observed EXIF and records the discrepancy.

Sources: dataset DOI and public Sycan Marsh subset.

Scope correction - 31 July 2026

FLAME 3 is retained as a detector-software side result, not evidence for candidate optics.

The replay reads native DJI M30T Celsius pixels, converts them to Kelvin, and passes them directly to compact-hot-v1. It does not apply an EmberScope optical prescription.

What the replay measures

Transfer of the current threshold detector to public DJI thermal scenes, including its measured no-fire clutter-rejection failure.

What it does not measure

Candidate focal length, F-number, throughput, PSF/MTF, image spread, vignetting, plate scale, GSD, detector sampling, NETD, or noise.

Current decision

No further FLAME 3 acquisition is required. RFS-069 closes without a burn-held-out acceptance claim, and the active project returns to optical candidates and prescriptions.

Possible future use

A separately authorized end-to-end simulation could transform scene radiometry through each candidate's measured optics and detector response before detection.

The 24 July approval selected the 100 mm F/1.8 custom-optics target and continued engineering work. The agent proposed the later FLAME 3 tasks while expanding that queue; they were not a separate detector-development request from Greg. Greg explicitly approved this scope and provenance correction on 31 July.

Validation principle

EmberScope should be tested as a radiometric survey payload, not as a generic fire-picture classifier.

A useful benchmark has to show detection of weak or small hot targets at the drone GSD and dwell time, rejection of rural no-fire clutter, and survival of held-out burns, sites, days, and sensor paths.

The plan keeps RGB-only, thermal-only, RGB/thermal, and RGB/radiometric-TIFF scores separate so visible smoke or flame cues cannot disguise weak thermal performance.

Validation stack

The headline split must hold out burns and backgrounds, not just shuffled frames.

Validation layer What it tests Required output
Burn-held-out public data Generalization across complete fire events rather than adjacent video frames. Precision, recall, specificity, sensitivity, F1, ROC/PR data, and confusion matrix by modality.
No-fire stress set False alarms from sun-heated clutter, vehicles, people, structures, smoke, shadows, water, and residual heat. False positives per flight minute and per surveyed hectare, with representative rejected examples.
Radiometric TIFF / raw thermal path Whether the detector chain works from temperature-like data rather than palette color alone. Thermal-only, RGB/thermal, and RGB/radiometric-TIFF results reported separately.
Small-hot-target simulation EmberScope's centimetre-scale target after GSD, dwell, blur, noise, and calibration error are known. Detection curves versus target size, target radiance or temperature, background, altitude, and threshold.
Local EmberScope field data Same detector, calibration kit, optics, geotagging, and survey profile as the payload under test. Field surrogate results and failed-case packet ready for engineering review.

False-positive policy

The negative set has to look like rural fire-service operating terrain, not a clean lab background.

Hot clutter

Sun-heated rocks, bare ground, roads, rooftops, metal gates, vehicles, and machinery.

Warm non-fire objects

People, livestock, buildings, camp equipment, and other warm objects that operators must not chase as fire.

Atmospheric and optical confusion

Smoke, dust, haze, shadows, clouds, water, reflective surfaces, and RGB/thermal alignment offsets.

Operational negatives

Pre-burn, post-burn, and same-terrain no-fire flights at comparable altitude, speed, time of day, and solar loading.

Acceptance gate

No detection claim should ship without provenance, held-out data, and failed examples.

A credible first report needs a burn-held-out public score, a no-fire stress score, a radiometric-TIFF or raw-radiometry score, an EmberScope-specific small-hot-target simulation, and a packet of missed detections and false positives for review.

These are detector-validation gates, not optical-design gates. They apply only if a future detector-validation effort is separately authorized.

Every run should record dataset source, version, checksum, split definition, model or rule version, threshold, detector assumptions, calibration inputs, GSD, altitude, frame rate, and reviewer label provenance. Raw datasets and bulky experiment outputs should stay outside routine public materials unless intentionally packaged for review.