Full-dataset experiment plan

Tracker for video-then-audio runs on the expanded corpus (764 segments, video-grouped splits). New runs write under outputs/; the earlier 290-segment runs live in outputs_subset/. Hardware: CPU. Primary metric: validation macro-F1 on remaining classes after label drops.

Decisions locked in

  • Imbalance: class-weighted cross-entropy (same recipe as prior baseline). Balanced mini-batch sampling is a later A/B — not stacked with weights.
  • Drop from train + eval (runtime filter in code; do not delete data/frames):
    • Function: Discussion (0 val)
    • Engagement: Half, Some (0 val)
    • Organization: keep all 3 (Small Group has val support)
  • Splits: keep current video-grouped train/val/test (541 / 88 / 135). No re-stratify for this pass.
  • Gate: ~≥0.03 val macro-F1 gain over the previous step on the same target, or clear minority F1 recovery without tanking majors.

Dataset context

Manifest: 764 segments · 73 linked videos · frames present for 763. After drops, Function is 5-class; Engagement is 3-class (All / Most / Not Oservable); Organization stays 3-class.

Video sequence (run in order)

V1 → V2 → V3 for Function, Organization, and Engagement. Stop after V1 to review metrics before V2. V4 only if V1–V3 show no meaningful gain.

ID Status What Details Output
V1 done Expanded baseline Frozen ResNet-18, 8 frames, mean-pool (--temporal-head none), weighted CE, 25 epochs, exclude zero-val labels.
Results (val / test macro-F1): Function 0.277 / 0.161 (5-class, dropped Discussion); Organization 0.307 / 0.308; Engagement 0.376 / 0.357 (3-class, dropped Half/Some).
outputs/{function,organization,engagement}/baseline/
V2 done Temporal head Same as V1 with --temporal-head transformer. Results (val / test): Function 0.259 / 0.199 (no gain vs V1); Organization 0.407 / 0.401 (+0.100); Engagement 0.434 / 0.404 (+0.057). outputs/.../temporal_head_8f/
V3 planned More frames Only if V2 helps or ties: re-extract --n-frames 16, retrain temporal-head at 16 frames outputs/.../temporal_head_16f/
V4 deferred VideoMAE / heavier video model Only if V1–V3 lack meaningful val macro-F1 gain outputs/.../videomae/ (TBD)

Audio sequence

Native-timing log-mel (mel_time / approach D) for all three targets — same structure as the video runs, with spectrogram features concatenated. Compare each run to that target’s best video-only score (Function V1; Org/Eng V2). ASR confidence / diarization may be added later as small cached side-channels on top of mel; not in this pass.

Video trunk per target

  • Function: V1 mean-pool (--temporal-head none) — temporal head did not help.
  • Organization: V2 transformer temporal head.
  • Engagement: V2 transformer temporal head.

Multimodal trainer must support --exclude-labels and --temporal-head (same as train_function.py) before these runs; wire that when executing A1–A3.

ID Status What Details Output
A0 done Cache mel_time Cached native-timing log-mel for 763 segments (extract_audio_features.py --sets mel_time). data/audio_features/mel_time/
A1 done Function + mel_time Video V1 mean-pool + native log-mel; weighted CE; Discussion dropped. Results: val 0.307 / test 0.238 (vs video V1 0.277 / 0.161) — +0.030 val, meaningful gain. outputs/function/mm_mel_time/
A1b done Function + mel + ASR uncertainty Same as A1 plus Whisper-tiny uncertainty stats (8-D; no transcript text). Results: val 0.321 / test 0.372 (vs A1 0.307 / 0.238) — +0.014 val; best Function so far. outputs/function/mm_mel_time_asr/
A2 done Engagement + mel_time Video V2 temporal head + native log-mel; Half/Some dropped. Results: val 0.320 / test 0.391 (vs video V2 0.434 / 0.404) — −0.114 val, audio hurt. outputs/engagement/mm_mel_time/
A2b done Engagement + mel + ASR Same as A2 plus ASR uncertainty. Results: val 0.304 / test 0.353 — still below video V2 (0.434); keep video-only for Engagement. outputs/engagement/mm_mel_time_asr/
A3 done Organization + mel_time Video V2 temporal head + native log-mel. Results: val 0.410 / test 0.324 (vs video V2 0.407 / 0.401) — essentially tied on val (+0.003); test worse. outputs/organization/mm_mel_time/
A3b done Organization + mel + ASR Same as A3 plus ASR uncertainty. Results: val 0.478 / test 0.401 (vs video V2 0.407) — +0.071 val; best Organization so far. outputs/organization/mm_mel_time_asr/

Later (not this pass)

  • Diarization summary stats as additive features (optional next)
  • Warped mel128, prosody-only, Whisper transcript embeddings, TF-IDF comments

Not in this plan (yet)

  • Default subset “v2” recipe (focal + unfreeze layer3/4) as the first full-data run
  • mel128, prosody, PANNs-only, Whisper transcripts, TF-IDF comments, broad ablation grid
  • ASR confidence / diarization (planned as post–A1–A3 add-ons)
  • Balanced sampling + unweighted CE (optional A/B after mel runs)
  • Rebuilding splits / stratified k-fold
  • Deleting rare-class frame folders from disk

How to run

V1 — baseline (all three targets)

python scripts/train_function.py --label-column function --loss ce --fine-tune-blocks 0 --dropout 0.3 --epochs 25 --n-frames 8 --temporal-head none --exclude-labels Discussion --output-dir outputs/function/baseline

python scripts/train_function.py --label-column organization --loss ce --fine-tune-blocks 0 --dropout 0.3 --epochs 25 --n-frames 8 --temporal-head none --output-dir outputs/organization/baseline

python scripts/train_function.py --label-column engagement --loss ce --fine-tune-blocks 0 --dropout 0.3 --epochs 25 --n-frames 8 --temporal-head none --exclude-labels Half Some --output-dir outputs/engagement/baseline

V2 — temporal head (8 frames)

python scripts/train_function.py --label-column function --loss ce --fine-tune-blocks 0 --dropout 0.3 --epochs 25 --n-frames 8 --temporal-head transformer --exclude-labels Discussion --output-dir outputs/function/temporal_head_8f
# same pattern for organization / engagement (engagement: --exclude-labels Half Some)

V3 — 16 frames (after re-extract)

python scripts/extract_segments.py --n-frames 16
python scripts/train_function.py --label-column function --loss ce --fine-tune-blocks 0 --dropout 0.3 --epochs 25 --n-frames 16 --temporal-head transformer --exclude-labels Discussion --output-dir outputs/function/temporal_head_16f

A0 — cache native-timing log-mel

# If WAVs missing:
python scripts/extract_audio.py --skip-existing

python scripts/extract_audio_features.py --feature-set mel_time --skip-existing

A1b — Function mel + ASR uncertainty

python scripts/extract_audio_features.py --sets asr_uncertainty --skip-existing
python scripts/train_multimodal_v2.py --audio-mode mel_time_asr --label-column function --loss ce --fine-tune-blocks 0 --temporal-head none --exclude-labels Discussion --output-dir outputs/function/mm_mel_time_asr

A1–A3 — multimodal (after wiring --exclude-labels / --temporal-head into multimodal trainer)

# A1 Function — V1 video trunk + mel_time
python scripts/train_multimodal_v2.py --audio-mode mel_time --label-column function --loss ce --fine-tune-blocks 0 --temporal-head none --exclude-labels Discussion --output-dir outputs/function/mm_mel_time

# A2 Engagement — V2 video trunk + mel_time
python scripts/train_multimodal_v2.py --audio-mode mel_time --label-column engagement --loss ce --fine-tune-blocks 0 --temporal-head transformer --exclude-labels Half Some --output-dir outputs/engagement/mm_mel_time

# A3 Organization — V2 video trunk + mel_time
python scripts/train_multimodal_v2.py --audio-mode mel_time --label-column organization --loss ce --fine-tune-blocks 0 --temporal-head transformer --output-dir outputs/organization/mm_mel_time

Approval / execution notes

Video V1–V2 and audio A0–A3 + ASR add-ons (A1b/A2b/A3b) are done. Best Function: mel + ASR (0.321). Best Organization: mel + ASR (0.478). Best Engagement: video temporal head (0.434) — keep video-only. See PROJECT_LOG.html.

Prior subset: mel_time helped Function/Engagement val slightly; Organization preferred video-only — full-data Org audio still does not clearly beat video.