# Following the Light: exploratory pix2pixHD results on Truth Beam projector-camera pairs, with a fixed-context counterfactual comparison

*Results package p2pv2v_20260906.*

This selected results package includes discriminator results for the six static-condition models and the temporal-generator tests. Discriminator tests were also run on both temporal-condition models; their results, per-pair records and discriminator logs are intentionally withheld from this package. The public driver log records the removal of two temporal-discriminator summary lines. Included maps, environment records and pairs manifests are in metrics/deferred/; published logs are in logs/deferred/.

8 of 8 runs evaluated. This package reports exploratory results from eight pix2pixHD runs on two sessions. The holdout is the contiguous final 10 percent of each session: 600 d2 frames (5392 to 5991) and 375 v10 frames (3368 to 3742). These are held-out frames from trained-on sessions and geometry; unseen-session generalisation is untested, and serial frames are dependent. Under the canonical prepared split, no held-out filename appears in train_A or train_B. Exact exclusion from the terminated box's /home/ubuntu/pairs_local training copy is unverified because no pre-termination manifest survives. A replacement-box manifest proves only the replacement copy. Full training-input closure for the reported weights requires recovery of the original disk digest or retraining from manifested inputs. The terminated box's dependency environment is unrecorded unless independently recovered; new-box records apply only to new reruns. Full-frame fidelity is measured at 2048x1152. Teacher-forced and previous-real results use the 973 eligible frames (the first held-out frame of each session, which has no eligible predecessor within the held-out split, is excluded); the emission and training-mean baselines use all 975.

For non-temporal runs, G verification compares the residual correlation of G(E_t) with B_t against G(E_s) with B_t, where s is one seeded different frame from the same session. Residuals subtract a separate mean for each session, computed from 200 evenly strided training captures at 512x288 grey, which attenuates the shared background (drift, motion, exposure, noise and model error remain). Temporal G rows compare complete temporal conditions (E_t, E_t-1, B_t-1 all change with s), so they measure whole-condition specificity; the counterfactual section below reports an exploratory current-emission comparison, subject to its cross-execution provenance limits. Paired wins depend on the single chosen partner; pooled AUROC compares the full score populations. Read the paired wins first.

D results (v2, 6 September 2026): every discriminator was rescored on a fresh box (canonical prepared pairs inventoried in metrics/deferred/manifests; metrics/deferred/ENVIRONMENT.json records the final follow-up environment, while per-process dependency versions were not recorded; the initial and final dependency snapshots differ, and the final snapshot does not establish the versions loaded by every earlier evaluator process or by the original teacher-forced generation) over 400 targets drawn from the 973 eligible held-out frames (predecessor also held out) by a recorded derangement (seed 2026090610, numpy PCG64; 237 d2 and 163 v10 targets; every partner in the target's session and different from the target; every target, partner and score persisted). For the six static-condition models, the summary D column reports full-frame capture-swap AUROC with paired wins out of 400 ("cs v2": the target's emission scored against the partner's capture). The detail table reports their full-frame and centre-crop capture-swap and emission-swap results; in their emission-swap test, only E_t is swapped. The six static-condition discriminators separate matched from mismatched pairs (full-frame capture swap 0.69 to 0.97; emission swap within 0.01 of it). Discriminator results are reported for the six static-condition models only; the temporal-condition models are reported through their generator tests (teacher-forced rows and the counterfactual section), and the published copy of the box driver log (logs/deferred/deferred_public.log) omits the two lines that summarised their discriminator rescoring. These are held-out frames of trained-on sessions and the discriminator scores are uncalibrated. For the six static-condition models, the earlier D results, whose sampler let 92 of 400 negatives cross sessions, are superseded and retained only inside their per-run JSON for provenance. The baseline rows are trivial predictors scored on the same frames. The train-free statistic and the non-temporal G verification carry the positive correspondence claim. The counterfactual comparison supplies supporting exploratory evidence with the provenance limits described below. The eight main-run videos contain 39 every-25th frames; the two teacher-forced videos contain 37. Figures are rounded; the JSON keeps full precision, including per-frame fidelity, G, D v2 and counterfactual rows.

| run | variant | schedule (epochs) | frames | PSNR (all) | SSIM (all) | PSNR d2 / v10 | D-verifier v2 (capture swap, full frame): AUROC (wins/400) | G-verifier: paired wins (pooled AUROC) |
|---|---|---|---|---|---|---|---|---|
| tb_base_b4 | baseline, batch 4 | 20 + 5 | 975 | 17.01 | 0.575 | 17.60 / 16.07 | 0.965 (387/400) cs v2 | 957/975 (AUROC 0.654) |
| tb_base_seed2 | second independently initialised baseline (no seed recorded) | 20 + 5 | 975 | 16.37 | 0.534 | 16.67 / 15.89 | 0.943 (388/400) cs v2 | 946/975 (AUROC 0.781) |
| tb_base_b8 | baseline, batch 8 | 20 + 5 | 975 | 17.22 | 0.587 | 17.86 / 16.20 | 0.692 (285/400) cs v2 | 959/975 (AUROC 0.701) |
| tb_base_ngf64 | wider generator (ngf 64) | 12 + 3 | 975 | 16.83 | 0.632 | 17.50 / 15.75 | 0.745 (296/400) cs v2 | 735/975 (AUROC 0.530) |
| tb_base_hires | baseline, 1024 px training crops, batch 2 | 6 + 2 | 975 | 16.30 | 0.592 | 16.60 / 15.82 | 0.712 (287/400) cs v2 | 969/975 (AUROC 0.933) |
| tb_temporal_b4 | temporal, 9-channel input (E_t, E_t-1, B_t-1) | 12 + 3 | 975 | 14.09 | 0.383 | 14.15 / 13.99 | not reported (temporal condition) | 934/975 (AUROC 0.818) |
| tb_temporal_b4_tf | temporal, 9-channel input (E_t, E_t-1, B_t-1) (teacher-forced: real previous capture as input) | 12 + 3 | 973 | 22.59 | 0.821 | 23.23 / 21.57 | not reported (temporal condition) | 969/973 (AUROC 0.996) whole temporal condition |
| tb_temporal_hires | temporal, 1024 px training crops, batch 2 | 6 + 2 | 975 | 11.62 | 0.464 | 12.88 / 9.59 | not reported (temporal condition) | 890/975 (AUROC 0.788) |
| tb_temporal_hires_tf | temporal, 1024 px training crops, batch 2 (teacher-forced: real previous capture as input) | 6 + 2 | 973 | 25.96 | 0.858 | 26.03 / 25.85 | not reported (temporal condition) | 973/973 (AUROC 1.000) whole temporal condition |
| tb_finetune2025 | initialised from the 2025 epoch-370 G and compatible D layers; the three D output heads newly initialised | 15 + 5 | 975 | 18.22 | 0.626 | 19.08 / 16.84 | 0.768 (375/400) cs v2 | 951/975 (AUROC 0.597) |
| emission | baseline: emission frame as the prediction | n/a | 975 | 7.39 | 0.266 | 7.32 / 7.49 | n/a | n/a |
| prev_real | baseline: previous real capture as the prediction | n/a | 973 | 20.32 | 0.805 | 20.65 / 19.77 | n/a | n/a |
| mean_train | baseline: mean training capture | n/a | 975 | 21.87 | 0.863 | 22.47 / 20.90 | n/a | n/a |

## Discriminator-as-verifier detail, v2 (six static-condition models; 400 targets, recorded within-session derangement seed 2026090610; pooled AUROC and paired wins/400)

| run | condition | capture swap crop512 | capture swap full | emission swap crop512 | emission swap full |
|---|---|---|---|---|---|
| tb_base_b4 | emission | 0.802 (331/400) | 0.965 (387/400) | 0.806 (397/400) | 0.960 (400/400) |
| tb_base_seed2 | emission | 0.922 (366/400) | 0.943 (388/400) | 0.927 (400/400) | 0.940 (400/400) |
| tb_base_b8 | emission | 0.603 (266/400) | 0.692 (285/400) | 0.603 (375/400) | 0.689 (399/400) |
| tb_base_ngf64 | emission | 0.607 (254/400) | 0.745 (296/400) | 0.609 (364/400) | 0.744 (397/400) |
| tb_base_hires | emission | 0.534 (215/400) | 0.712 (287/400) | 0.531 (316/400) | 0.714 (398/400) |
| tb_finetune2025 | emission | 0.717 (306/400) | 0.768 (375/400) | 0.717 (400/400) | 0.764 (400/400) |

## Counterfactual comparison using retained teacher-forced outputs and new alternative-emission outputs (973 eligible frames; recorded within-session derangement seed 2026090611)

The matched arm reuses the morning teacher-forced outputs in latest_tf/fake; the alternative-emission arm was generated on the replacement box in latest_cf/fake. The intended comparison changes E_t to the recorded E_s while retaining E_t-1 and the real B_t-1. The new loader changes only the current emission. Both output sets are scored against B_t using full correlation and residual correlation, with a separate mean for each session computed from 200 evenly strided training captures at 512x288 grey.

The previous-real predictor ties by construction because its output is identical in both arms. The last two columns report mean residual correlation with B_t-1; their similar rounded means do not rule out per-frame persistence effects or other confounds.

The checkpoint path was reused, but the record does not bind both generation executions to identical generator checkpoint bytes, history-input bytes or runtime versions. These results are exploratory evidence consistent with current-emission correspondence under the intended fixed context; differences between the two generation executions remain unresolved. They concern held-out frames of trained-on sessions and do not establish unseen-session generalisation or validated liveness.

| run | frames | residual corr: matched wins/frames (pooled AUROC) | full corr: wins (AUROC) | d2 / v10 residual AUROC | matched output: residual corr with B_t-1 | counterfactual output: residual corr with B_t-1 |
|---|---|---|---|---|---|---|
| tb_temporal_b4 | 973 | 972/973 (AUROC 0.950) | 973/973 (0.967) | 0.986 / 0.925 | 0.374 | 0.374 |
| tb_temporal_hires | 973 | 973/973 (AUROC 0.998) | 973/973 (0.998) | 1.000 / 0.995 | 0.420 | 0.420 |

## Generator-as-verifier detail (paired wins / frames, pooled AUROC; 512x288 grey; residual PSNR omitted because subtracting the same mean from both images leaves it identical to full PSNR)

| run | frames | residual corr | full corr | full PSNR |
|---|---|---|---|---|
| tb_base_b4 | 975 | 957/975 (AUROC 0.654) | 842/975 (AUROC 0.823) | 820/975 (AUROC 0.577) |
| tb_base_seed2 | 975 | 946/975 (AUROC 0.781) | 851/975 (AUROC 0.850) | 827/975 (AUROC 0.678) |
| tb_base_b8 | 975 | 959/975 (AUROC 0.701) | 881/975 (AUROC 0.883) | 864/975 (AUROC 0.594) |
| tb_base_ngf64 | 975 | 735/975 (AUROC 0.530) | 644/975 (AUROC 0.655) | 627/975 (AUROC 0.511) |
| tb_base_hires | 975 | 969/975 (AUROC 0.933) | 967/975 (AUROC 0.965) | 955/975 (AUROC 0.807) |
| tb_temporal_b4 | 975 | 934/975 (AUROC 0.818) | 767/975 (AUROC 0.787) | 661/975 (AUROC 0.668) |
| tb_temporal_b4_tf | 973 | 969/973 (AUROC 0.996) | 971/973 (AUROC 0.995) | 970/973 (AUROC 0.983) |
| tb_temporal_hires | 975 | 890/975 (AUROC 0.788) | 779/975 (AUROC 0.691) | 612/975 (AUROC 0.592) |
| tb_temporal_hires_tf | 973 | 973/973 (AUROC 1.000) | 973/973 (AUROC 1.000) | 973/973 (AUROC 0.999) |
| tb_finetune2025 | 975 | 951/975 (AUROC 0.597) | 849/975 (AUROC 0.762) | 851/975 (AUROC 0.538) |

## Train-free coupling statistic (grid Pearson correlation, emission vs capture)

Each mismatch uses a different held-out capture from the same session. Each grid used a different seeded shuffle. This is a random-negative correspondence test, not a quality score or validated liveness test.

| grid | pooled AUROC | matched > mismatched (paired) | d2 AUROC | v10 AUROC |
|---|---|---|---|---|
| 4 | 0.683 | 900 / 975 | 0.688 | 0.677 |
| 8 | 0.756 | 902 / 975 | 0.768 | 0.736 |
| 16 | 0.720 | 832 / 975 | 0.726 | 0.710 |
| 32 | 0.700 | 820 / 975 | 0.709 | 0.687 |
| 64 | 0.695 | 781 / 975 | 0.689 | 0.704 |
