Back to projects

ML tooling / Real-time systems / Engineering notes

ML Training Inspector

I built a live training dashboard and verified its charts and epoch table against a reproducible synthetic PyTorch CPU run.

Inspect the repository

Reviewed 6 October 2026 · Local verification recorded 4 October · Implementation and notes published

Published verification record
ML Training Inspector showing measured synthetic CPU demo loss, accuracy and gradient charts
Actual dashboard from a seeded two-epoch synthetic CPU run. The observed metrics verify the visualization pipeline, not real-world predictive quality. Checkpoint resume is unavailable. Select the image for the full-size capture.

The problem

A training loop can produce numbers without making its behavior understandable. I built a browser dashboard to expose loss, accuracy, gradients, class-level results, and run state while training progresses.

My contribution

I connected PyTorch training to FastAPI and WebSockets, built React charts and an accessible epoch table, and added manual stopping and checkpoint saving. Recent work corrects example-weighted loss aggregation, bounds per-client queues, handles disconnects and terminal outcomes, and adds a seeded, download-free CPU demonstration.

One decision: keep run ownership explicit

The backend deliberately supports one shared run in one worker. Conflicting starts receive HTTP 409. Bounded client queues protect memory, and reconnects recover status and the latest epoch metadata rather than promising missed chart history. Stopped and error states remain distinct from successful completion.

Verification and evidence scope

The published record reports 15 backend tests and two frontend tests, plus a successful production build. A real browser CPU run rendered charts and an epoch table whose values matched the demo. The seeded SimpleCNN run used two epochs, 160 synthetic training examples, and 80 validation examples. Model/optimizer loading reproduced validation metrics, but scheduler, random-state, and sampler state are not preserved, so resume is absent. The measured synthetic accuracy is not evidence of generalization.

Demo

Follow the published README to install the CPU demo dependencies and run python scripts/demo.py. For the browser walkthrough, start the backend with one worker and the frontend, choose the synthetic dataset, and run two SimpleCNN epochs. Inspect the charts and matching epoch table, then explore stop and snapshot controls. The capture on this page comes from the verified synthetic run. No public hosted service is linked.

Limits and next work

GPU runs, full CIFAR-10 training, ResNet9 training, containers, and cross-platform repeatability were not verified in that record. There is no authentication or per-user ownership. Heuristic gradient signals have no detection-quality evaluation; some near-zero gradients can be benign. Checkpoint resume and reconnect chart-history replay are not implemented.

Get in touch