Multimodal classroom segment classification — full-dataset report

We retrain on the expanded corpus (764 segments) for Function, Organization, and Engagement, with native-timing log-mel and ASR uncertainty add-ons. Companion logs: PROJECT_LOG.html, EXPERIMENT_PLAN.html.

1. Three prediction tasks

Each annotated segment carries three independent labels. We trained separate classifiers for Function (instructional activity), Organization (whole class / groups / individual), and Engagement (how many students appear engaged). The same video-grouped splits and segment clips were reused; only the target column changed.

2. Data and inputs

Label counts (full manifest, n=764)

CategoryClassCount%
FunctionDirection22028.8
Working18524.2
Lecture15420.2
Transition10313.5
IRE/MS9512.4
Discussion70.9
OrganizationWhole Class54571.3
Individual Work16221.2
Small Group577.5
EngagementAll48062.8
Most16922.1
Not Oservable8811.5
Some202.6
Half70.9

Prior subset was the same spreadsheet restricted to 29 on-disk videos (Discussion entirely missing). Macro-F1 there was noisier (val 59 / test 41).

3. Evaluation metric

Models are scored with macro-averaged F1 (macro-F1): for each class we compute precision and recall, take their harmonic mean (F1), then average across classes with equal weight. Unlike accuracy, macro-F1 does not let the model look good by always predicting the majority class.

We select checkpoints by validation macro-F1 (88 segments after drops). Test (128–135 segments) is reported but not used for selection.

Training loss: class-weighted cross-entropy (weight ∝ 1 / train frequency). That encourages rare-label learning during optimization; macro-F1 measures held-out balance afterward.

4. Modeling approaches (full-data sequence)

Unlike the subset report’s A–D grid (including warped mel128 and focal fine-tune “v2”), the full-data plan starts from a frozen ResNet baseline and adds temporal pooling and audio only where they help.

V1 — Video baseline

  • Video: ResNet-18 (ImageNet); backbone frozen; 8 frames → mean pool → linear head.
  • Loss: weighted CE; 25 epochs; AdamW on the head only.

Reference “frames only.” Best Function video trunk (temporal head did not help Function).

V2 — Temporal head (8 frames)

  • Video: same frozen ResNet-18; replace mean-pool with a 1-layer TransformerEncoder (4 heads) over the 8 frame embeddings, then mean over time → linear head.
  • Loss: weighted CE; 25 epochs.

Helped Organization (+0.10) and Engagement (+0.06); did not help Function.

A1–A3 — Native-timing log-mel multimodal

  • Video trunk: Function = V1 mean-pool; Org/Eng = V2 temporal head.
  • Audio: cached native-timing log-mel (mel_time) → 1D CNN + mask → 128-D; concat with video → classifier.
  • Loss: weighted CE; frozen backbone; 25 epochs.

Mel helped Function; hurt Engagement; Organization roughly tied with video on val.

A1b–A3b — Mel + ASR uncertainty

  • Same as mel multimodal, plus an 8-D Whisper-tiny uncertainty vector (no transcript text / no MiniLM embedding).
  • Features include avg log-prob, no-speech prob, compression ratio, segment count, speech fraction, log-prob std, char rate, empty-transcript flag.

Best Function and best Organization so far; Engagement still prefers video-only.

Not carried forward from the prior report

  • Warped mel128 (fixed 64×128 resize) — underperformed native mel_time on Function before; not re-run.
  • Focal loss + unfreezing layer3/4 as the default first recipe — reserved; full-data V1 stayed frozen + CE.
  • Prosody-only, PANNs-only, Whisper transcript embeddings, TF-IDF comments — deferred / previously weak.

5. Audio encoding

5.1 From WAV to log-mel

  1. Load segment WAV; resample to 16 kHz mono if needed.
  2. Mel spectrogram: 64 mel bins, FFT 400, hop 160 (~10 ms per frame).
  3. Log compression: log(mel + 1e-9). Networks always see log-mel.

5.2 Native timing (mel_time)

Keep one time column per hop (~100 frames/s). Clips shorter than ~5.1 s are zero-padded to 512 frames with a mask; longer clips are truncated to the first 512 frames (~5 s of audio content). A 1D conv + masked mean pool yields a 128-D vector.

5.3 ASR uncertainty (not transcripts)

On the subset, Whisper → MiniLM transcript embeddings were unstable: ASR often emits fluent text on noisy classroom audio. Here we keep only confidence / uncertainty statistics from whisper-tiny and concatenate them with mel (and video). The transcript string is never fed to the classifier.

6. Results (macro-F1)

Validation = model selection; test = held-out videos. Bold / green rows are the current best model per target.

Function (5-class; Discussion dropped)

ChoiceValTest
V1 baseline — frozen ResNet-18, 8 frames, mean-pool, weighted CE 0.2770.161
V2 temporal head — transformer over 8 frame embeddings 0.2590.199
A1 mel_time — V1 video + native log-mel 0.3070.238
A1b mel+ASR — V1 video + log-mel + Whisper uncertainty (no transcript) 0.3210.372

Organization (3-class)

ChoiceValTest
V1 baseline — frozen ResNet-18, mean-pool 8f 0.3070.308
V2 temporal head — transformer over frames 0.4070.401
A3 mel_time — V2 video + native log-mel 0.4100.324
A3b mel+ASR — V2 video + log-mel + ASR uncertainty 0.4780.401

Engagement (3-class; Half/Some dropped)

ChoiceValTest
V1 baseline — frozen ResNet-18, mean-pool 8f 0.3760.357
V2 temporal head — transformer over frames 0.4340.404
A2 mel_time — V2 video + log-mel 0.3200.391
A2b mel+ASR — V2 video + log-mel + ASR uncertainty 0.3040.353

Summary

  • Function: best = mel + ASR uncertainty (val 0.321 / test 0.372).
  • Organization: best = mel + ASR (val 0.478 / test 0.401); large gain over video-only.
  • Engagement: best = video temporal head (val 0.434 / test 0.404); mel and ASR hurt.

7. Function examples (best val model: mel + ASR)

Validation clips from the best Function model (mm_mel_time_asr, n=88 after dropping Discussion). Thumbnail is the first of eight sampled frames (click to enlarge). Each row includes a playable WAV of the segment. Comments are annotations only, not model inputs.

Correctly classified (5)

Frame Segment + audio Comment True Predicted
1910 · 0:14:00–0:16:29
Teacher telling students what they needed to to Direction Direction
1908 · 0:00:00–0:09:22
Teacher talking about the lesson - Teacher projecting his laptop on the screen - Lecture Lecture
1909 · 0:03:31–0:04:32
Students going to grab iPads Transition Transition
1909 · 0:08:24–0:32:00
Students working on their tasks - Teacher moving around and helping students Working Working
1908 · 0:18:49–0:19:44
Teacher telling students to use the Wash brush to erase the shape they created - Teacher showing a new shape on the s... Direction Direction

Misclassified (5)

Frame Segment + audio Comment True Predicted
1909 · 0:04:32–0:08:24
Teacher giving directions to students Direction Lecture
1924 · 0:04:26–0:11:44
teacher asking questions IRE/MS Lecture
1921 · 0:10:00–0:26:34
teacher explaining the lesson Lecture Working
1909 · 0:32:00–0:36:43
Students leaving the class after finishing the tasks Transition Working
1908 · 0:16:34–0:18:49
Students working on their laptops - teacher moving around and talking to students Working Lecture

8. Findings

9. Future directions (estimated usefulness)

  1. Diarization summary features — cached speaker-count / overlap / speech-fraction stats. Org and Function already benefit from audio clarity cues; “who talks how much” is a natural next side-channel for grouping.
  2. V3: 16 frames + temporal head — denser frame sampling on the video-led task. Engagement rejected audio; more visual temporal detail is the cleanest next upgrade there.
  3. Stronger video model (VideoMAE / TimeSformer) — pretrained clip transformer instead of bag-of-ResNet frames. Largest architectural jump if denser frames plateau, especially for Engagement and Function transitions.
  4. Light backbone fine-tune (layer4) on current Function/Org winners. A little adaptation of ImageNet filters may help with ~500 train clips without full 3D cost.
  5. Balanced mini-batch sampling (vs weighted CE only). Minorities (Transition, Small Group, Most) still drag macro-F1; this is the deferred imbalance A/B.
  6. 3D CNN (R3D / R(2+1)D) — true spatiotemporal convolutions on the frame stack. More video-like than pooling embeddings, but heavy on CPU and easy to overfit.
  7. Larger ASR for uncertainty only (e.g. Whisper-small). Same confidence idea with better estimates if tiny-Whisper signal is real but noisy.
  8. Larger audio encoders (PANNs / BEATs) fused with mel — rich pretrained audio beside log-mel. Mixed on the subset; lower priority now that mel+ASR already works for Fun/Org.
  9. Stratified video re-split / GroupKFold — better coverage of rares and uncertainty on metrics. Cannot invent Discussion/Half examples; mainly reporting hygiene.