Compute-Optimal Data Selection for Video-Generation Model Adaptation
This is a 16-week Research Engineer flagship investigating whether deliberate video-data selection can improve adaptation quality per GPU-hour over uniform sampling. The reference model is LTX-Video 2B with parameter-efficient adaptation. The project is not a claim of frontier-model pretraining and does not use a custom toy DiT as the primary result.
At a fixed training-token, update-step, seed, prompt-suite, and compute budget, which data-selection strategy produces the best quality–diversity–cost trade-off when adapting a video diffusion transformer?
- Quality-only filtering improves visual fidelity but reduces motion and semantic coverage.
- Motion/semantic stratification improves temporal behavior and tail-category coverage over uniform sampling.
- A quality–diversity Pareto sampler gives the strongest overall result per GPU-hour.
- Data choice remains a material contributor after controlling the model, optimizer, steps, resolution, clip length, and seed.
| ID | Baseline | Why it exists | Required evidence |
|---|---|---|---|
| E0 | Official LTXV 2B LoRA reproduction | Validate the trainer, resume path, memory use, and cost on the actual cloud setup | 100–200 steps, resumed checkpoint, generated sample, run manifest |
| B0 | Unadapted zero-shot model | Establish pre-adaptation quality and detect regressions | Frozen prompt suite, versioned outputs, metric report |
| B1 | Uniform-sampled adaptation | Primary research control for all data-selection claims | Equal tokens/steps, two seeds for the confirmatory comparison |
E0 is an engineering reproduction, not the research control. B0 is the quality reference. B1 is the experimental control.
- acquire and version an OpenVid-1M subset;
- validate clips and compute technical, motion, semantic, and duplicate features;
- implement uniform, quality-only, stratified, and Pareto selection policies;
- reproduce the official LTXV adaptation path before modifying it;
- run controlled training experiments and two-seed confirmation;
- evaluate visual quality, text alignment, temporal consistency, diversity, failure slices, and cost;
- measure single- and two-GPU behavior without claiming cluster-scale expertise;
- profile and improve the inference path based on measured bottlenecks;
- package an observable service, load test it, inject failures, and verify rollback;
- publish reproducible artifacts, negative results, and an honest technical report.
A task is not complete because code ran. It is complete only when it has a versioned artifact such as a manifest, test, experiment card, raw metric file, profiler trace, benchmark report, deployment log, or pull-request link. Predictions must be written before comparison runs. Failed runs and wrong assumptions belong in ENGINEERING_LOG.md.
parallax-video-adaptation/
├── planning/ # Formula-driven 16-week Excel tracker
├── configs/
│ ├── data/ # acquisition, validation, feature and sampling configs
│ ├── training/ # E0/B1/treatment training configs
│ ├── experiments/ # frozen controlled-comparison configs
│ └── serving/ # service and load-test configuration
├── data/ # manifests and small samples only; raw media stays external
├── src/parallax_video/
│ ├── data/ # ingest, validation, features, deduplication, sampling, sharding
│ ├── training/ # official-trainer integration, controls, checkpointing, distributed runs
│ ├── evaluation/ # metrics, prompt suite, ablations and failure slicing
│ └── systems/ # profiling, inference, serving, observability and cost accounting
├── scripts/ # thin reproducible CLI entrypoints
├── experiments/
│ ├── cards/ # preregistered predictions and decision rules
│ ├── registry/ # run index and artifact locations
│ └── results/ # machine-readable result summaries
├── evaluation/ # prompts, human-eval material and generated reports
├── benchmarks/ # training, inference and service benchmark outputs
├── deployment/ # API, load tests and observability assets
├── tests/ # unit, integration and bounded distributed tests
├── docs/ # research specification, architecture and reproducibility
├── reports/ # paper/report figures and tables
- Open
planning/Video_Generative_RE_16_Week_Execution_Tracker.xlsxand read Start Here, Baselines & Gates, and Week 1. - Read
docs/research_spec.mdand freeze the initial claims before collecting results. - Create the dataset card and manifest under
data/manifests/. - Complete E0, then B0, then B1. Do not begin treatment claims before all three gates pass.
- Create an experiment card from
experiments/cards/TEMPLATE.mdbefore every comparison.
- The archived custom DiT/RingAttention/Triton exercises are learning history, not evidence for the flagship result.
- A two-GPU validation demonstrates correctness and measurement discipline, not frontier-cluster experience.
- A deployed demo is production-oriented, not production-proven without real users and sustained traffic.
- Results apply to the documented dataset, model, resolution, clip length, prompts, and compute envelope.