Executive summary (current)
Best model so far (by val macro-F1, per target)
- Function: A1b
mm_mel_time_asr— val 0.321 / test 0.372. - Organization: A3b
mm_mel_time_asr— val 0.478 / test 0.401 (beats video V2 0.407). - Engagement: V2
temporal_head_8f— val 0.434 / test 0.404 (mel+ASR 0.304 still hurts).
Setup notes
- Class-weighted CE; frozen ResNet-18; 8 frames; native-timing log-mel for audio runs.
- Dropped zero-val labels via
--exclude-labels. - Next optional: diarization; V3 16 frames for Engagement video.
Results tables (macro-F1)
Function
Excluded: Discussion. Classes: Direction, IRE/MS, Lecture, Transition, Working.
| Run | Architecture | Val F1 | Test F1 | n train/val/test | Outcome |
|---|---|---|---|---|---|
baseline (V1) |
ResNet-18 frozen + mean-pool 8f; weighted CE | 0.277 | 0.161 | 534 / 88 / 134 | Full-data video baseline |
temporal_head_8f (V2) |
ResNet-18 frozen + transformer temporal head; 8f | 0.259 | 0.199 | 534 / 88 / 134 | No meaningful val gain vs V1 (−0.018) |
temporal_head_16f (V3) |
Same + 16 frames | — | — | — | Deferred |
mm_mel_time (A1) |
V1 video + native log-mel | 0.307 | 0.238 | 534 / 88 / 134 | +0.030 vs V1 |
mm_mel_time_asr (A1b) |
V1 video + log-mel + Whisper ASR uncertainty (8-D) | 0.321 | 0.372 | 534 / 88 / 134 | +0.014 vs A1 — Function leader |
Organization
All 3 classes kept (Small Group has val support).
| Run | Architecture | Val F1 | Test F1 | n train/val/test | Outcome |
|---|---|---|---|---|---|
baseline (V1) |
ResNet-18 frozen + mean-pool 8f; weighted CE | 0.307 | 0.308 | 540 / 88 / 135 | Below prior subset Org baseline (~0.42) |
temporal_head_8f (V2) |
ResNet-18 frozen + transformer temporal head; 8f | 0.407 | 0.401 | 540 / 88 / 135 | +0.100 val vs V1 |
temporal_head_16f (V3) |
Same + 16 frames | — | — | — | Deferred |
mm_mel_time (A3) |
V2 video + native log-mel | 0.410 | 0.324 | 540 / 88 / 135 | Tied val vs V2; test worse |
mm_mel_time_asr (A3b) |
V2 video + log-mel + ASR uncertainty | 0.478 | 0.401 | 540 / 88 / 135 | +0.071 vs V2 — Org leader |
Engagement
Excluded: Half, Some. Classes: All, Most, Not Oservable.
| Run | Architecture | Val F1 | Test F1 | n train/val/test | Outcome |
|---|---|---|---|---|---|
baseline (V1) |
ResNet-18 frozen + mean-pool 8f; weighted CE | 0.376 | 0.357 | 520 / 88 / 128 | Strong V1; Most / Not Oservable still weak |
temporal_head_8f (V2) |
ResNet-18 frozen + transformer temporal head; 8f | 0.434 | 0.404 | 520 / 88 / 128 | +0.057 val vs V1 — preferred Eng model |
temporal_head_16f (V3) |
Same + 16 frames | — | — | — | Deferred |
mm_mel_time (A2) |
V2 video + native log-mel | 0.320 | 0.391 | 520 / 88 / 128 | −0.114 vs V2 — audio hurt |
mm_mel_time_asr (A2b) |
V2 video + log-mel + ASR uncertainty | 0.304 | 0.353 | 520 / 88 / 128 | Still worse than video V2 |
V1 per-class F1 (validation)
Function
| Class | Val F1 | Support |
|---|---|---|
| Direction | 0.258 | 22 |
| IRE/MS | 0.242 | 13 |
| Lecture | 0.390 | 18 |
| Transition | 0.095 | 9 |
| Working | 0.400 | 26 |
Organization
| Class | Val F1 | Support |
|---|---|---|
| Individual Work | 0.093 | 18 |
| Small Group | 0.167 | 8 |
| Whole Class | 0.661 | 62 |
Engagement
| Class | Val F1 | Support |
|---|---|---|
| All | 0.800 | 66 |
| Most | 0.129 | 13 |
| Not Oservable | 0.200 | 9 |
Best model detail
Strongest full-data checkpoint by val macro-F1: Engagement V2 temporal head (0.434). Function still led by V1 baseline (0.277). Confusion matrices / sample grids can be expanded later.
Confusion matrix (Function V1, validation, n=88)
Rows = true label. Columns = predicted. Darker blue = more segments.
| Pred: Direction | Pred: IRE/MS | Pred: Lecture | Pred: Transition | Pred: Working | |
|---|---|---|---|---|---|
| True: Direction | 4 | 6 | 3 | 3 | 6 |
| True: IRE/MS | 2 | 4 | 5 | 0 | 2 |
| True: Lecture | 1 | 2 | 8 | 2 | 5 |
| True: Transition | 1 | 3 | 3 | 1 | 1 |
| True: Working | 1 | 5 | 4 | 6 | 10 |
Sample inputs: true vs predicted
Data scope
| Source | Segments | Videos | Notes |
|---|---|---|---|
Full segments.csv / manifest | 764 | 73 | Local edu_all/ |
| Train / val / test (by video) | 541 / 88 / 135 | — | Before label drops; 1 segment missing frames |
| Function after drop Discussion | 534 / 88 / 134 | — | 5 classes |
| Engagement after drop Half/Some | 520 / 88 / 128 | — | 3 classes |
| Organization (no drop) | 540 / 88 / 135 | — | 3 classes |
split Whole-video splits, seed 42. No segment leakage.
End-to-end pipeline
Inputs & outputs per modality
| Modality | Input | Representation | Used in training |
|---|---|---|---|
| video | 8 frames / segment (16 later) | 224×224 ImageNet-normalized RGB; mean-pool or temporal head | V1–V3 |
| audio mel_time | Cached log-mel | Conv1d + masked pool | A1+ (not started) |
Chronological log
Phase 0 — Expanded data ready
data 764 segments, 73 videos, frames for 763. Plan published in EXPERIMENT_PLAN.html.
Phase V1 — Video baseline (full data)
model Frozen ResNet-18, weighted CE, 8 frames, mean-pool. Drop Discussion / Half / Some.
Results: Function 0.277 / 0.161 · Organization 0.307 / 0.308 · Engagement 0.376 / 0.357 (val / test).
outputs/*/baseline/Phase V2 — Temporal head (8 frames)
model Same as V1 with --temporal-head transformer.
Results: Function 0.259 / 0.199 · Organization 0.407 / 0.401 · Engagement 0.434 / 0.404 (val / test).
Gate: Org + Eng clear gains; Function no gain → keep V1 for Function trunk.
outputs/*/temporal_head_8f/Phase A0–A3 — Native log-mel multimodal
audio Cached mel_time for 763 segments; multimodal with best video trunk per target.
Results: Function 0.307 / 0.238 (+mel helps) · Engagement 0.320 / 0.391 (hurts) · Organization 0.410 / 0.324 (tied val).
outputs/*/mm_mel_time/Phase A1b — Function mel + ASR uncertainty
audio Whisper-tiny confidence stats (8-D; no transcript embedding) concatenated with mel_time.
Result: Function val 0.321 / test 0.372 — best Function so far (+0.014 vs A1).
outputs/function/mm_mel_time_asr/Phase A2b / A3b — Eng & Org mel + ASR
audio Same ASR add-on on V2 video trunks.
Results: Engagement 0.304 / 0.353 (still below video 0.434). Organization 0.478 / 0.401 (+0.071 vs video V2) — new Org leader.
outputs/engagement/mm_mel_time_asr/ outputs/organization/mm_mel_time_asr/Model experiments (video)
Same numbers as the per-target tables above; kept here for parity with the subset project log.
| Run ID | Target | Architecture | Val F1 | Test F1 | Outcome |
|---|---|---|---|---|---|
outputs/function/baseline |
Function | ResNet-18 frozen; mean-pool 8f; weighted CE | 0.277 | 0.161 | V1 Function leader |
outputs/function/temporal_head_8f |
Function | Transformer temporal head; 8f | 0.259 | 0.199 | No meaningful val gain |
outputs/organization/baseline |
Organization | ResNet-18 frozen; mean-pool 8f; weighted CE | 0.307 | 0.308 | V1 Org |
outputs/organization/temporal_head_8f |
Organization | Transformer temporal head; 8f | 0.407 | 0.401 | V2 Org leader (+0.100) |
outputs/engagement/baseline |
Engagement | ResNet-18 frozen; mean-pool 8f; weighted CE | 0.376 | 0.357 | V1 Engagement |
outputs/engagement/temporal_head_8f |
Engagement | Transformer temporal head; 8f | 0.434 | 0.404 | V2 Engagement leader (+0.057) |
Audio representation runs
A0–A3 complete. Native-timing log-mel only (no ASR/diarization yet).
| Run | Target | Audio feature | Val F1 | Test F1 | Notes |
|---|---|---|---|---|---|
mm_mel_time (A1) |
Function | mel_time | 0.307 | 0.238 | +0.030 vs video V1 0.277 |
mm_mel_time_asr (A1b) |
Function | mel_time + ASR uncertainty | 0.321 | 0.372 | +0.014 vs A1; no transcript text used |
mm_mel_time (A2) |
Engagement | mel_time | 0.320 | 0.391 | −0.114 vs video V2 0.434 |
mm_mel_time_asr (A2b) |
Engagement | mel_time + ASR uncertainty | 0.304 | 0.353 | Still below video V2 0.434 |
mm_mel_time (A3) |
Organization | mel_time | 0.410 | 0.324 | ~tied vs video V2 0.407; test drop |
mm_mel_time_asr (A3b) |
Organization | mel_time + ASR uncertainty | 0.478 | 0.401 | +0.071 vs video V2 — Org leader |
Lessons so far
- Dropping zero-val classes makes Engagement metrics readable (3-way) vs prior 5-way collapse on rares.
- Temporal head (V2) helped Organization (+0.10) and Engagement (+0.06) but not Function (−0.02).
- Native log-mel helped Function (+0.03) but hurt Engagement (−0.11); Organization essentially tied — keep video-only for Org/Eng.
- ASR uncertainty helped Function (+0.014) and Organization (+0.071 vs video); Engagement still prefers video-only.
- Organization V2 video (0.407) recovered near the old subset baseline; mel+ASR (0.478) goes further.
- Function Transition remains weak; mel + ASR helps overall Function macro-F1 modestly.
Artifacts
| Path | Contents |
|---|---|
| outputs/function/baseline/ | V1 Function metrics, checkpoint, val CM CSV |
| outputs/organization/baseline/ | V1 Organization |
| outputs/engagement/baseline/ | V1 Engagement |
| outputs/function/mm_mel_time/ | A1 Function + mel |
| outputs/function/mm_mel_time_asr/ | A1b Function + mel + ASR (best Function) |
| outputs/engagement/mm_mel_time/ | A2 Engagement + mel (worse than video) |
| outputs/engagement/mm_mel_time_asr/ | A2b Engagement + mel + ASR (still below video) |
| outputs/organization/mm_mel_time/ | A3 Organization + mel (tied val) |
| outputs/organization/mm_mel_time_asr/ | A3b Organization + mel + ASR (best Org) |
| data/audio_features/mel_time/ | A0 cached native log-mel (763) |
| data/audio_features/asr_uncertainty/ | Whisper-tiny uncertainty vectors (763) |
| docs/EXPERIMENT_PLAN.html | Run sequence and commands |
| docs/PROJECT_LOG.html | Prior 290-segment subset log |