1. Three prediction tasks
Each annotated segment carries three independent labels. We trained separate classifiers for Function (instructional activity), Organization (whole class / groups / individual), and Engagement (how many students appear engaged). The same video-grouped splits and segment clips were reused; only the target column changed.
2. Data and inputs
- 764 segments from 73 videos under
edu_all/(vs prior local subset: 290 / 29). - Per segment: 8 RGB frames (224×224, ImageNet normalization) + mono WAV at 16 kHz.
- Split by video: train 541 / validation 88 / test 135 (no segment leakage). One segment lacks frames.
- Written comments are excluded from all reported models (not available at deployment).
-
Label drops (runtime filter; files kept on disk): Function drops
Discussion(n=7, 0 in val); Engagement dropsHalfandSome(0 in val). Organization keeps all 3 classes. These labels had no (or effectively no) validation examples, so keeping them would let train-only classes inflate or distort macro-F1 without a fair held-out check — a form of evaluation leakage. We drop them at train/eval time instead.
Label counts (full manifest, n=764)
| Category | Class | Count | % |
|---|---|---|---|
| Function | Direction | 220 | 28.8 |
| Working | 185 | 24.2 | |
| Lecture | 154 | 20.2 | |
| Transition | 103 | 13.5 | |
| IRE/MS | 95 | 12.4 | |
| Discussion | 7 | 0.9 | |
| Organization | Whole Class | 545 | 71.3 |
| Individual Work | 162 | 21.2 | |
| Small Group | 57 | 7.5 | |
| Engagement | All | 480 | 62.8 |
| Most | 169 | 22.1 | |
| Not Oservable | 88 | 11.5 | |
| Some | 20 | 2.6 | |
| Half | 7 | 0.9 |
Prior subset was the same spreadsheet restricted to 29 on-disk videos (Discussion entirely missing). Macro-F1 there was noisier (val 59 / test 41).
3. Evaluation metric
Models are scored with macro-averaged F1 (macro-F1): for each class we compute precision and recall, take their harmonic mean (F1), then average across classes with equal weight. Unlike accuracy, macro-F1 does not let the model look good by always predicting the majority class.
We select checkpoints by validation macro-F1 (88 segments after drops). Test (128–135 segments) is reported but not used for selection.
Training loss: class-weighted cross-entropy (weight ∝ 1 / train frequency). That encourages rare-label learning during optimization; macro-F1 measures held-out balance afterward.
4. Modeling approaches (full-data sequence)
Unlike the subset report’s A–D grid (including warped mel128 and focal fine-tune “v2”), the full-data plan starts from a frozen ResNet baseline and adds temporal pooling and audio only where they help.
V1 — Video baseline
- Video: ResNet-18 (ImageNet); backbone frozen; 8 frames → mean pool → linear head.
- Loss: weighted CE; 25 epochs; AdamW on the head only.
Reference “frames only.” Best Function video trunk (temporal head did not help Function).
V2 — Temporal head (8 frames)
- Video: same frozen ResNet-18; replace mean-pool with a 1-layer TransformerEncoder (4 heads) over the 8 frame embeddings, then mean over time → linear head.
- Loss: weighted CE; 25 epochs.
Helped Organization (+0.10) and Engagement (+0.06); did not help Function.
A1–A3 — Native-timing log-mel multimodal
- Video trunk: Function = V1 mean-pool; Org/Eng = V2 temporal head.
- Audio: cached native-timing log-mel (
mel_time) → 1D CNN + mask → 128-D; concat with video → classifier. - Loss: weighted CE; frozen backbone; 25 epochs.
Mel helped Function; hurt Engagement; Organization roughly tied with video on val.
A1b–A3b — Mel + ASR uncertainty
- Same as mel multimodal, plus an 8-D Whisper-tiny uncertainty vector (no transcript text / no MiniLM embedding).
- Features include avg log-prob, no-speech prob, compression ratio, segment count, speech fraction, log-prob std, char rate, empty-transcript flag.
Best Function and best Organization so far; Engagement still prefers video-only.
Not carried forward from the prior report
- Warped mel128 (fixed 64×128 resize) — underperformed native mel_time on Function before; not re-run.
- Focal loss + unfreezing layer3/4 as the default first recipe — reserved; full-data V1 stayed frozen + CE.
- Prosody-only, PANNs-only, Whisper transcript embeddings, TF-IDF comments — deferred / previously weak.
5. Audio encoding
5.1 From WAV to log-mel
- Load segment WAV; resample to 16 kHz mono if needed.
- Mel spectrogram: 64 mel bins, FFT 400, hop 160 (~10 ms per frame).
- Log compression:
log(mel + 1e-9). Networks always see log-mel.
5.2 Native timing (mel_time)
Keep one time column per hop (~100 frames/s). Clips shorter than ~5.1 s are zero-padded to 512 frames with a mask; longer clips are truncated to the first 512 frames (~5 s of audio content). A 1D conv + masked mean pool yields a 128-D vector.
5.3 ASR uncertainty (not transcripts)
On the subset, Whisper → MiniLM transcript embeddings were unstable: ASR often emits fluent text on noisy classroom audio. Here we keep only confidence / uncertainty statistics from whisper-tiny and concatenate them with mel (and video). The transcript string is never fed to the classifier.
6. Results (macro-F1)
Validation = model selection; test = held-out videos. Bold / green rows are the current best model per target.
Function (5-class; Discussion dropped)
| Choice | Val | Test |
|---|---|---|
| V1 baseline — frozen ResNet-18, 8 frames, mean-pool, weighted CE | 0.277 | 0.161 |
| V2 temporal head — transformer over 8 frame embeddings | 0.259 | 0.199 |
| A1 mel_time — V1 video + native log-mel | 0.307 | 0.238 |
| A1b mel+ASR — V1 video + log-mel + Whisper uncertainty (no transcript) | 0.321 | 0.372 |
Organization (3-class)
| Choice | Val | Test |
|---|---|---|
| V1 baseline — frozen ResNet-18, mean-pool 8f | 0.307 | 0.308 |
| V2 temporal head — transformer over frames | 0.407 | 0.401 |
| A3 mel_time — V2 video + native log-mel | 0.410 | 0.324 |
| A3b mel+ASR — V2 video + log-mel + ASR uncertainty | 0.478 | 0.401 |
Engagement (3-class; Half/Some dropped)
| Choice | Val | Test |
|---|---|---|
| V1 baseline — frozen ResNet-18, mean-pool 8f | 0.376 | 0.357 |
| V2 temporal head — transformer over frames | 0.434 | 0.404 |
| A2 mel_time — V2 video + log-mel | 0.320 | 0.391 |
| A2b mel+ASR — V2 video + log-mel + ASR uncertainty | 0.304 | 0.353 |
Summary
- Function: best = mel + ASR uncertainty (val 0.321 / test 0.372).
- Organization: best = mel + ASR (val 0.478 / test 0.401); large gain over video-only.
- Engagement: best = video temporal head (val 0.434 / test 0.404); mel and ASR hurt.
7. Function examples (best val model: mel + ASR)
Validation clips from the best Function model
(mm_mel_time_asr, n=88 after dropping Discussion).
Thumbnail is the first of eight sampled frames (click to enlarge).
Each row includes a playable WAV of the segment. Comments are annotations only, not model inputs.
Correctly classified (5)
| Frame | Segment + audio | Comment | True | Predicted |
|---|---|---|---|---|
|
1910 · 0:14:00–0:16:29 |
Teacher telling students what they needed to to | Direction | Direction | |
|
1908 · 0:00:00–0:09:22 |
Teacher talking about the lesson - Teacher projecting his laptop on the screen - | Lecture | Lecture | |
|
1909 · 0:03:31–0:04:32 |
Students going to grab iPads | Transition | Transition | |
|
1909 · 0:08:24–0:32:00 |
Students working on their tasks - Teacher moving around and helping students | Working | Working | |
|
1908 · 0:18:49–0:19:44 |
Teacher telling students to use the Wash brush to erase the shape they created - Teacher showing a new shape on the s... | Direction | Direction |
Misclassified (5)
| Frame | Segment + audio | Comment | True | Predicted |
|---|---|---|---|---|
|
1909 · 0:04:32–0:08:24 |
Teacher giving directions to students | Direction | Lecture | |
|
1924 · 0:04:26–0:11:44 |
teacher asking questions | IRE/MS | Lecture | |
|
1921 · 0:10:00–0:26:34 |
teacher explaining the lesson | Lecture | Working | |
|
1909 · 0:32:00–0:36:43 |
Students leaving the class after finishing the tasks | Transition | Working | |
|
1908 · 0:16:34–0:18:49 |
Students working on their laptops - teacher moving around and talking to students | Working | Lecture |
8. Findings
- Expanding from 290→764 segments makes metrics more stable; label drops remove zero-val rares from the primary score.
- A temporal head over frame embeddings helps Organization and Engagement; Function prefers plain mean-pool + audio.
- Native-timing log-mel helps Function; alone it does not clearly help Org and hurts Engagement.
- ASR uncertainty (not transcript text) is a useful additive signal for Function and especially Organization.
- Engagement remains primarily a visual / attention problem on this data — stacking classroom audio degraded val macro-F1.
9. Future directions (estimated usefulness)
- Diarization summary features — cached speaker-count / overlap / speech-fraction stats. Org and Function already benefit from audio clarity cues; “who talks how much” is a natural next side-channel for grouping.
- V3: 16 frames + temporal head — denser frame sampling on the video-led task. Engagement rejected audio; more visual temporal detail is the cleanest next upgrade there.
- Stronger video model (VideoMAE / TimeSformer) — pretrained clip transformer instead of bag-of-ResNet frames. Largest architectural jump if denser frames plateau, especially for Engagement and Function transitions.
- Light backbone fine-tune (layer4) on current Function/Org winners. A little adaptation of ImageNet filters may help with ~500 train clips without full 3D cost.
- Balanced mini-batch sampling (vs weighted CE only). Minorities (Transition, Small Group, Most) still drag macro-F1; this is the deferred imbalance A/B.
- 3D CNN (R3D / R(2+1)D) — true spatiotemporal convolutions on the frame stack. More video-like than pooling embeddings, but heavy on CPU and easy to overfit.
- Larger ASR for uncertainty only (e.g. Whisper-small). Same confidence idea with better estimates if tiny-Whisper signal is real but noisy.
- Larger audio encoders (PANNs / BEATs) fused with mel — rich pretrained audio beside log-mel. Mixed on the subset; lower priority now that mel+ASR already works for Fun/Org.
- Stratified video re-split / GroupKFold — better coverage of rares and uncertainty on metrics. Cannot invent Discussion/Half examples; mainly reporting hygiene.