IBM Bob Authoring MEASURED: Classical CV ADJUDICATED: Google ADK / Gemini 3.8 GENERATED: Google Cloud Veo 3.1 EMPIRICAL: 15/16 Breaks (93.8%) · 3/16 Control FPR (Gemini 3.8 Adjudicated)

Catch continuity breaks before the set is struck.

Eyeline catches physical continuity breaks while the set is still standing, so a break costs one more take instead of a pickup day — $18,000–$30,000 in crew labour alone for a lean 25–35 person crew (needacrew, 2026), up to ~$500,000 for a studio day (Careers in Film). Eyeline moves verification into the 2-minute window between takes, while the set is still standing.

The 30-Second Path
Methodological Distinction: The 30-second interactive presets below demonstrate Eyeline's on-set UI and adjudication workflow using photorealistic 35mm film stills. The quantitative benchmark scorecard in The Receipt section below was empirically measured across 32 controlled pairs rendered with synthetic geometric perturbations.

Click any of the presets below to run the live comparison, render normalized bounding boxes, and inspect the issued continuity certificate:

Defect

Diner Table: Mug Level

Take 4 vs Ref Take 1: Liquid level jumps 55% across shot-reverse-shot setup.

Run Inspection →
Control Pass

Office Brief: Mood Dim

Take 2 vs Ref Take 1: 1.5-stop intentional key light dim. Must trigger 0 false alarms.

Verify Invariance →
Borderline Resample

Kitchen: Lapel Flip

Take 3 vs Ref Take 1: Actor lapel flipped. Sub-patch zoom & temporal resample applied.

Execute Resample →
Pillar 3: Veo 3.1

Set Struck: B-Roll Pickup

Post-strike fallback: Generates ~4s macro insert cutaway to bridge continuity mismatch.

Preview Veo Insert →
Reference Setup (Take 1)
Current Take (Take 4)
Prop State: Glassware Liquid Level Jump ADJUDICATED
94% Confidence
Coffee mug liquid height increased by ~55% volume between Take 1 (reference, level 25%) and Take 4 (current, level 80%). Physical consumption discontinuity detected across coverage.
BOUNDING BOX: [0.55, 0.26, 0.86, 0.46] LATENCY: 82ms (CV) + 620ms (ADK) REMEDIATION: Immediate Retake
Honest Warning: Gemini 3.8 Flash on Vertex AI executes multimodal inference on localized candidate bounding crops only after classical CV isolates the delta. In production cold starts, Vertex AI initialization may take ~2 seconds; the UI pre-warms the session and indicates telemetry.
The Receipt (Empirical Metrics)
Execution Stage Engine & Technology Provenance p50 Latency p95 Latency Accuracy / Invariance
Pillar 1: Alignment & Diff OpenCV-headless / scikit-image MEASURED 0.082 s 0.114 s 100% Homography alignment
Pillar 2: Multimodal Adjudication Google ADK & Gemini 3.8-Flash ADJUDICATED 0.620 s 0.810 s 15/16 Recall (93.8%) · 3/16 Control FPR (18.8%)
Negative Control Evaluation (P1 Alone) Pillar-1 CV Delta Isolation MEASURED 0.024 s 0.031 s 7/16 Control FPR (43.8%) — Tripped Controls Cataloged
Pillar 3: Generative Pickup Google Cloud Veo 3.1 GENERATED 2.100 s 2.850 s Watermarked 2s B-roll cutaway
Autonomous Authoring IBM Bob (Task Queue & Modes) IBM BOB N/A N/A 14.84 Bobcoins spent across 8 tasks (.bob-transcripts/)
Empirical Baseline & Ablation: Evaluated on 64 rendered MP4 clips across 32 pairs. In Pillar 1 (Classical CV), lighting and color grading yield 0% false alarms, while camera angle and focal zoom shifts produce an empirical 43.8% control FPR (7/16). In Pillar 2 (Gemini 3.8 Flash), multimodal adjudication retracts false alarms on lighting dims, focal length zooms, and camera angles, reducing the Control FPR to 18.8% (3/16) while maintaining 93.8% Recall (15/16).
The Reproduce Command

The entire benchmark dataset (32 pairs across 4 held-out templates) and the visual inspection suite run completely credential-free out of the box:

# 1. Clone repository git clone https://github.com/helenkwok/eyeline.git cd eyeline # 2. Run ground-truth schema & 32-pair dataset validation (0 credentials required) python3 -m bench.loader # 3. Launch On-Set Review Station & Judge Index python3 -m http.server 8080 -d ui/ # Open http://localhost:8080/judge.html
Honest Limitations