Decisions locked in
- Imbalance: class-weighted cross-entropy (same recipe as prior baseline). Balanced mini-batch sampling is a later A/B — not stacked with weights.
- Drop from train + eval (runtime filter in code; do not delete
data/frames):- Function:
Discussion(0 val) - Engagement:
Half,Some(0 val) - Organization: keep all 3 (
Small Grouphas val support)
- Function:
- Splits: keep current video-grouped train/val/test (541 / 88 / 135). No re-stratify for this pass.
- Gate: ~≥0.03 val macro-F1 gain over the previous step on the same target, or clear minority F1 recovery without tanking majors.
Dataset context
Manifest: 764 segments · 73 linked videos · frames present for 763. After drops, Function is 5-class; Engagement is 3-class (All / Most / Not Oservable); Organization stays 3-class.
Video sequence (run in order)
V1 → V2 → V3 for Function, Organization, and Engagement. Stop after V1 to review metrics before V2. V4 only if V1–V3 show no meaningful gain.
| ID | Status | What | Details | Output |
|---|---|---|---|---|
| V1 | done | Expanded baseline | Frozen ResNet-18, 8 frames, mean-pool (--temporal-head none), weighted CE, 25 epochs, exclude zero-val labels.Results (val / test macro-F1): Function 0.277 / 0.161 (5-class, dropped Discussion); Organization 0.307 / 0.308; Engagement 0.376 / 0.357 (3-class, dropped Half/Some). |
outputs/{function,organization,engagement}/baseline/ |
| V2 | done | Temporal head | Same as V1 with --temporal-head transformer.
Results (val / test):
Function 0.259 / 0.199 (no gain vs V1);
Organization 0.407 / 0.401 (+0.100);
Engagement 0.434 / 0.404 (+0.057).
|
outputs/.../temporal_head_8f/ |
| V3 | planned | More frames | Only if V2 helps or ties: re-extract --n-frames 16, retrain temporal-head at 16 frames |
outputs/.../temporal_head_16f/ |
| V4 | deferred | VideoMAE / heavier video model | Only if V1–V3 lack meaningful val macro-F1 gain | outputs/.../videomae/ (TBD) |
Audio sequence
Native-timing log-mel (mel_time / approach D) for all three targets —
same structure as the video runs, with spectrogram features concatenated.
Compare each run to that target’s best video-only score (Function V1; Org/Eng V2).
ASR confidence / diarization may be added later as small cached side-channels on top of mel; not in this pass.
Video trunk per target
- Function: V1 mean-pool (
--temporal-head none) — temporal head did not help. - Organization: V2 transformer temporal head.
- Engagement: V2 transformer temporal head.
Multimodal trainer must support --exclude-labels and --temporal-head
(same as train_function.py) before these runs; wire that when executing A1–A3.
| ID | Status | What | Details | Output |
|---|---|---|---|---|
| A0 | done | Cache mel_time |
Cached native-timing log-mel for 763 segments
(extract_audio_features.py --sets mel_time).
|
data/audio_features/mel_time/ |
| A1 | done | Function + mel_time | Video V1 mean-pool + native log-mel; weighted CE; Discussion dropped. Results: val 0.307 / test 0.238 (vs video V1 0.277 / 0.161) — +0.030 val, meaningful gain. | outputs/function/mm_mel_time/ |
| A1b | done | Function + mel + ASR uncertainty | Same as A1 plus Whisper-tiny uncertainty stats (8-D; no transcript text). Results: val 0.321 / test 0.372 (vs A1 0.307 / 0.238) — +0.014 val; best Function so far. | outputs/function/mm_mel_time_asr/ |
| A2 | done | Engagement + mel_time | Video V2 temporal head + native log-mel; Half/Some dropped. Results: val 0.320 / test 0.391 (vs video V2 0.434 / 0.404) — −0.114 val, audio hurt. | outputs/engagement/mm_mel_time/ |
| A2b | done | Engagement + mel + ASR | Same as A2 plus ASR uncertainty. Results: val 0.304 / test 0.353 — still below video V2 (0.434); keep video-only for Engagement. | outputs/engagement/mm_mel_time_asr/ |
| A3 | done | Organization + mel_time | Video V2 temporal head + native log-mel. Results: val 0.410 / test 0.324 (vs video V2 0.407 / 0.401) — essentially tied on val (+0.003); test worse. | outputs/organization/mm_mel_time/ |
| A3b | done | Organization + mel + ASR | Same as A3 plus ASR uncertainty. Results: val 0.478 / test 0.401 (vs video V2 0.407) — +0.071 val; best Organization so far. | outputs/organization/mm_mel_time_asr/ |
Later (not this pass)
- Diarization summary stats as additive features (optional next)
- Warped mel128, prosody-only, Whisper transcript embeddings, TF-IDF comments
Not in this plan (yet)
- Default subset “v2” recipe (focal + unfreeze layer3/4) as the first full-data run
- mel128, prosody, PANNs-only, Whisper transcripts, TF-IDF comments, broad ablation grid
- ASR confidence / diarization (planned as post–A1–A3 add-ons)
- Balanced sampling + unweighted CE (optional A/B after mel runs)
- Rebuilding splits / stratified k-fold
- Deleting rare-class frame folders from disk
How to run
V1 — baseline (all three targets)
python scripts/train_function.py --label-column function --loss ce --fine-tune-blocks 0 --dropout 0.3 --epochs 25 --n-frames 8 --temporal-head none --exclude-labels Discussion --output-dir outputs/function/baseline python scripts/train_function.py --label-column organization --loss ce --fine-tune-blocks 0 --dropout 0.3 --epochs 25 --n-frames 8 --temporal-head none --output-dir outputs/organization/baseline python scripts/train_function.py --label-column engagement --loss ce --fine-tune-blocks 0 --dropout 0.3 --epochs 25 --n-frames 8 --temporal-head none --exclude-labels Half Some --output-dir outputs/engagement/baseline
V2 — temporal head (8 frames)
python scripts/train_function.py --label-column function --loss ce --fine-tune-blocks 0 --dropout 0.3 --epochs 25 --n-frames 8 --temporal-head transformer --exclude-labels Discussion --output-dir outputs/function/temporal_head_8f # same pattern for organization / engagement (engagement: --exclude-labels Half Some)
V3 — 16 frames (after re-extract)
python scripts/extract_segments.py --n-frames 16 python scripts/train_function.py --label-column function --loss ce --fine-tune-blocks 0 --dropout 0.3 --epochs 25 --n-frames 16 --temporal-head transformer --exclude-labels Discussion --output-dir outputs/function/temporal_head_16f
A0 — cache native-timing log-mel
# If WAVs missing: python scripts/extract_audio.py --skip-existing python scripts/extract_audio_features.py --feature-set mel_time --skip-existing
A1b — Function mel + ASR uncertainty
python scripts/extract_audio_features.py --sets asr_uncertainty --skip-existing python scripts/train_multimodal_v2.py --audio-mode mel_time_asr --label-column function --loss ce --fine-tune-blocks 0 --temporal-head none --exclude-labels Discussion --output-dir outputs/function/mm_mel_time_asr
A1–A3 — multimodal (after wiring --exclude-labels / --temporal-head into multimodal trainer)
# A1 Function — V1 video trunk + mel_time python scripts/train_multimodal_v2.py --audio-mode mel_time --label-column function --loss ce --fine-tune-blocks 0 --temporal-head none --exclude-labels Discussion --output-dir outputs/function/mm_mel_time # A2 Engagement — V2 video trunk + mel_time python scripts/train_multimodal_v2.py --audio-mode mel_time --label-column engagement --loss ce --fine-tune-blocks 0 --temporal-head transformer --exclude-labels Half Some --output-dir outputs/engagement/mm_mel_time # A3 Organization — V2 video trunk + mel_time python scripts/train_multimodal_v2.py --audio-mode mel_time --label-column organization --loss ce --fine-tune-blocks 0 --temporal-head transformer --output-dir outputs/organization/mm_mel_time
Approval / execution notes
Video V1–V2 and audio A0–A3 + ASR add-ons (A1b/A2b/A3b) are done. Best Function: mel + ASR (0.321). Best Organization: mel + ASR (0.478). Best Engagement: video temporal head (0.434) — keep video-only. See PROJECT_LOG.html.
Prior subset: mel_time helped Function/Engagement val slightly; Organization preferred video-only — full-data Org audio still does not clearly beat video.