Classroom Video Classification — Full-dataset Project Log

Running record for experiments on the expanded corpus (edu_all, 764 segments). Companion to the 290-segment subset log (docs_subset/PROJECT_LOG.html). Plan: EXPERIMENT_PLAN.html. Collaborator report: COLLABORATOR_REPORT.html. Metric: macro-F1. Checkpoint selection: validation (88 segments after drops). Outputs under outputs/.

Executive summary (current)

Best model so far (by val macro-F1, per target)

  • Function: A1b mm_mel_time_asr — val 0.321 / test 0.372.
  • Organization: A3b mm_mel_time_asr — val 0.478 / test 0.401 (beats video V2 0.407).
  • Engagement: V2 temporal_head_8f — val 0.434 / test 0.404 (mel+ASR 0.304 still hurts).
outputs/function/mm_mel_time_asr/ outputs/organization/mm_mel_time_asr/ outputs/engagement/temporal_head_8f/

Setup notes

  • Class-weighted CE; frozen ResNet-18; 8 frames; native-timing log-mel for audio runs.
  • Dropped zero-val labels via --exclude-labels.
  • Next optional: diarization; V3 16 frames for Engagement video.

Results tables (macro-F1)

Function

Excluded: Discussion. Classes: Direction, IRE/MS, Lecture, Transition, Working.

Run Architecture Val F1 Test F1 n train/val/test Outcome
baseline (V1) ResNet-18 frozen + mean-pool 8f; weighted CE 0.277 0.161 534 / 88 / 134 Full-data video baseline
temporal_head_8f (V2) ResNet-18 frozen + transformer temporal head; 8f 0.259 0.199 534 / 88 / 134 No meaningful val gain vs V1 (−0.018)
temporal_head_16f (V3) Same + 16 frames — — — Deferred
mm_mel_time (A1) V1 video + native log-mel 0.307 0.238 534 / 88 / 134 +0.030 vs V1
mm_mel_time_asr (A1b) V1 video + log-mel + Whisper ASR uncertainty (8-D) 0.321 0.372 534 / 88 / 134 +0.014 vs A1 — Function leader

Organization

All 3 classes kept (Small Group has val support).

Run Architecture Val F1 Test F1 n train/val/test Outcome
baseline (V1) ResNet-18 frozen + mean-pool 8f; weighted CE 0.307 0.308 540 / 88 / 135 Below prior subset Org baseline (~0.42)
temporal_head_8f (V2) ResNet-18 frozen + transformer temporal head; 8f 0.407 0.401 540 / 88 / 135 +0.100 val vs V1
temporal_head_16f (V3) Same + 16 frames — — — Deferred
mm_mel_time (A3) V2 video + native log-mel 0.410 0.324 540 / 88 / 135 Tied val vs V2; test worse
mm_mel_time_asr (A3b) V2 video + log-mel + ASR uncertainty 0.478 0.401 540 / 88 / 135 +0.071 vs V2 — Org leader

Engagement

Excluded: Half, Some. Classes: All, Most, Not Oservable.

Run Architecture Val F1 Test F1 n train/val/test Outcome
baseline (V1) ResNet-18 frozen + mean-pool 8f; weighted CE 0.376 0.357 520 / 88 / 128 Strong V1; Most / Not Oservable still weak
temporal_head_8f (V2) ResNet-18 frozen + transformer temporal head; 8f 0.434 0.404 520 / 88 / 128 +0.057 val vs V1 — preferred Eng model
temporal_head_16f (V3) Same + 16 frames — — — Deferred
mm_mel_time (A2) V2 video + native log-mel 0.320 0.391 520 / 88 / 128 −0.114 vs V2 — audio hurt
mm_mel_time_asr (A2b) V2 video + log-mel + ASR uncertainty 0.304 0.353 520 / 88 / 128 Still worse than video V2

V1 per-class F1 (validation)

Function

ClassVal F1Support
Direction0.25822
IRE/MS0.24213
Lecture0.39018
Transition0.0959
Working0.40026

Organization

ClassVal F1Support
Individual Work0.09318
Small Group0.1678
Whole Class0.66162

Engagement

ClassVal F1Support
All0.80066
Most0.12913
Not Oservable0.2009

Best model detail

Strongest full-data checkpoint by val macro-F1: Engagement V2 temporal head (0.434). Function still led by V1 baseline (0.277). Confusion matrices / sample grids can be expanded later.

Confusion matrix (Function V1, validation, n=88)

Rows = true label. Columns = predicted. Darker blue = more segments.

Low count → high count
Pred: Direction Pred: IRE/MS Pred: Lecture Pred: Transition Pred: Working
True: Direction 4 6 3 3 6
True: IRE/MS 2 4 5 0 2
True: Lecture 1 2 8 2 5
True: Transition 1 3 3 1 1
True: Working 1 5 4 6 10

Sample inputs: true vs predicted

Not filled yet — export example frames / predictions after more runs settle.

Data scope

SourceSegmentsVideosNotes
Full segments.csv / manifest76473Local edu_all/
Train / val / test (by video)541 / 88 / 135—Before label drops; 1 segment missing frames
Function after drop Discussion534 / 88 / 134—5 classes
Engagement after drop Half/Some520 / 88 / 128—3 classes
Organization (no drop)540 / 88 / 135—3 classes

split Whole-video splits, seed 42. No segment leakage.

End-to-end pipeline

segments.csv + edu_all/ → prepare_dataset.py (link, QC, splits, manifest) → extract_segments.py → data/frames/ (8 JPEGs/segment) → train_function.py (--exclude-labels, --temporal-head) → outputs/{function,organization,engagement}/...

Inputs & outputs per modality

ModalityInputRepresentationUsed in training
video 8 frames / segment (16 later) 224×224 ImageNet-normalized RGB; mean-pool or temporal head V1–V3
audio mel_time Cached log-mel Conv1d + masked pool A1+ (not started)

Chronological log

Phase 0 — Expanded data ready

data 764 segments, 73 videos, frames for 763. Plan published in EXPERIMENT_PLAN.html.

Phase V1 — Video baseline (full data)

model Frozen ResNet-18, weighted CE, 8 frames, mean-pool. Drop Discussion / Half / Some.

Results: Function 0.277 / 0.161 · Organization 0.307 / 0.308 · Engagement 0.376 / 0.357 (val / test).

outputs/*/baseline/

Phase V2 — Temporal head (8 frames)

model Same as V1 with --temporal-head transformer.

Results: Function 0.259 / 0.199 · Organization 0.407 / 0.401 · Engagement 0.434 / 0.404 (val / test).

Gate: Org + Eng clear gains; Function no gain → keep V1 for Function trunk.

outputs/*/temporal_head_8f/

Phase A0–A3 — Native log-mel multimodal

audio Cached mel_time for 763 segments; multimodal with best video trunk per target.

Results: Function 0.307 / 0.238 (+mel helps) · Engagement 0.320 / 0.391 (hurts) · Organization 0.410 / 0.324 (tied val).

outputs/*/mm_mel_time/

Phase A1b — Function mel + ASR uncertainty

audio Whisper-tiny confidence stats (8-D; no transcript embedding) concatenated with mel_time.

Result: Function val 0.321 / test 0.372 — best Function so far (+0.014 vs A1).

outputs/function/mm_mel_time_asr/

Phase A2b / A3b — Eng & Org mel + ASR

audio Same ASR add-on on V2 video trunks.

Results: Engagement 0.304 / 0.353 (still below video 0.434). Organization 0.478 / 0.401 (+0.071 vs video V2) — new Org leader.

outputs/engagement/mm_mel_time_asr/ outputs/organization/mm_mel_time_asr/

Model experiments (video)

Same numbers as the per-target tables above; kept here for parity with the subset project log.

Run ID Target Architecture Val F1 Test F1 Outcome
outputs/function/baseline Function ResNet-18 frozen; mean-pool 8f; weighted CE 0.277 0.161 V1 Function leader
outputs/function/temporal_head_8f Function Transformer temporal head; 8f 0.259 0.199 No meaningful val gain
outputs/organization/baseline Organization ResNet-18 frozen; mean-pool 8f; weighted CE 0.307 0.308 V1 Org
outputs/organization/temporal_head_8f Organization Transformer temporal head; 8f 0.407 0.401 V2 Org leader (+0.100)
outputs/engagement/baseline Engagement ResNet-18 frozen; mean-pool 8f; weighted CE 0.376 0.357 V1 Engagement
outputs/engagement/temporal_head_8f Engagement Transformer temporal head; 8f 0.434 0.404 V2 Engagement leader (+0.057)

Audio representation runs

A0–A3 complete. Native-timing log-mel only (no ASR/diarization yet).

Run Target Audio feature Val F1 Test F1 Notes
mm_mel_time (A1) Function mel_time 0.307 0.238 +0.030 vs video V1 0.277
mm_mel_time_asr (A1b) Function mel_time + ASR uncertainty 0.321 0.372 +0.014 vs A1; no transcript text used
mm_mel_time (A2) Engagement mel_time 0.320 0.391 −0.114 vs video V2 0.434
mm_mel_time_asr (A2b) Engagement mel_time + ASR uncertainty 0.304 0.353 Still below video V2 0.434
mm_mel_time (A3) Organization mel_time 0.410 0.324 ~tied vs video V2 0.407; test drop
mm_mel_time_asr (A3b) Organization mel_time + ASR uncertainty 0.478 0.401 +0.071 vs video V2 — Org leader

Lessons so far

Artifacts

PathContents
outputs/function/baseline/V1 Function metrics, checkpoint, val CM CSV
outputs/organization/baseline/V1 Organization
outputs/engagement/baseline/V1 Engagement
outputs/function/mm_mel_time/A1 Function + mel
outputs/function/mm_mel_time_asr/A1b Function + mel + ASR (best Function)
outputs/engagement/mm_mel_time/A2 Engagement + mel (worse than video)
outputs/engagement/mm_mel_time_asr/A2b Engagement + mel + ASR (still below video)
outputs/organization/mm_mel_time/A3 Organization + mel (tied val)
outputs/organization/mm_mel_time_asr/A3b Organization + mel + ASR (best Org)
data/audio_features/mel_time/A0 cached native log-mel (763)
data/audio_features/asr_uncertainty/Whisper-tiny uncertainty vectors (763)
docs/EXPERIMENT_PLAN.htmlRun sequence and commands
docs/PROJECT_LOG.htmlPrior 290-segment subset log