Skip to content

Repository files navigation

Parallax: Compute-Optimal Video Model Adaptation

Compute-Optimal Data Selection for Video-Generation Model Adaptation

This is a 16-week Research Engineer flagship investigating whether deliberate video-data selection can improve adaptation quality per GPU-hour over uniform sampling. The reference model is LTX-Video 2B with parameter-efficient adaptation. The project is not a claim of frontier-model pretraining and does not use a custom toy DiT as the primary result.

Research question

At a fixed training-token, update-step, seed, prompt-suite, and compute budget, which data-selection strategy produces the best quality–diversity–cost trade-off when adapting a video diffusion transformer?

Primary hypotheses

  1. Quality-only filtering improves visual fidelity but reduces motion and semantic coverage.
  2. Motion/semantic stratification improves temporal behavior and tail-category coverage over uniform sampling.
  3. A quality–diversity Pareto sampler gives the strongest overall result per GPU-hour.
  4. Data choice remains a material contributor after controlling the model, optimizer, steps, resolution, clip length, and seed.

Baseline contract

ID Baseline Why it exists Required evidence
E0 Official LTXV 2B LoRA reproduction Validate the trainer, resume path, memory use, and cost on the actual cloud setup 100–200 steps, resumed checkpoint, generated sample, run manifest
B0 Unadapted zero-shot model Establish pre-adaptation quality and detect regressions Frozen prompt suite, versioned outputs, metric report
B1 Uniform-sampled adaptation Primary research control for all data-selection claims Equal tokens/steps, two seeds for the confirmatory comparison

E0 is an engineering reproduction, not the research control. B0 is the quality reference. B1 is the experimental control.

Active scope

  • acquire and version an OpenVid-1M subset;
  • validate clips and compute technical, motion, semantic, and duplicate features;
  • implement uniform, quality-only, stratified, and Pareto selection policies;
  • reproduce the official LTXV adaptation path before modifying it;
  • run controlled training experiments and two-seed confirmation;
  • evaluate visual quality, text alignment, temporal consistency, diversity, failure slices, and cost;
  • measure single- and two-GPU behavior without claiming cluster-scale expertise;
  • profile and improve the inference path based on measured bottlenecks;
  • package an observable service, load test it, inject failures, and verify rollback;
  • publish reproducible artifacts, negative results, and an honest technical report.

Evidence rule

A task is not complete because code ran. It is complete only when it has a versioned artifact such as a manifest, test, experiment card, raw metric file, profiler trace, benchmark report, deployment log, or pull-request link. Predictions must be written before comparison runs. Failed runs and wrong assumptions belong in ENGINEERING_LOG.md.

Repository map

parallax-video-adaptation/
├── planning/                 # Formula-driven 16-week Excel tracker
├── configs/
│   ├── data/                 # acquisition, validation, feature and sampling configs
│   ├── training/             # E0/B1/treatment training configs
│   ├── experiments/          # frozen controlled-comparison configs
│   └── serving/              # service and load-test configuration
├── data/                     # manifests and small samples only; raw media stays external
├── src/parallax_video/
│   ├── data/                 # ingest, validation, features, deduplication, sampling, sharding
│   ├── training/             # official-trainer integration, controls, checkpointing, distributed runs
│   ├── evaluation/           # metrics, prompt suite, ablations and failure slicing
│   └── systems/              # profiling, inference, serving, observability and cost accounting
├── scripts/                  # thin reproducible CLI entrypoints
├── experiments/
│   ├── cards/                # preregistered predictions and decision rules
│   ├── registry/             # run index and artifact locations
│   └── results/              # machine-readable result summaries
├── evaluation/               # prompts, human-eval material and generated reports
├── benchmarks/               # training, inference and service benchmark outputs
├── deployment/               # API, load tests and observability assets
├── tests/                    # unit, integration and bounded distributed tests
├── docs/                     # research specification, architecture and reproducibility
├── reports/                  # paper/report figures and tables

Start here

  1. Open planning/Video_Generative_RE_16_Week_Execution_Tracker.xlsx and read Start Here, Baselines & Gates, and Week 1.
  2. Read docs/research_spec.md and freeze the initial claims before collecting results.
  3. Create the dataset card and manifest under data/manifests/.
  4. Complete E0, then B0, then B1. Do not begin treatment claims before all three gates pass.
  5. Create an experiment card from experiments/cards/TEMPLATE.md before every comparison.

Non-claims

  • The archived custom DiT/RingAttention/Triton exercises are learning history, not evidence for the flagship result.
  • A two-GPU validation demonstrates correctness and measurement discipline, not frontier-cluster experience.
  • A deployed demo is production-oriented, not production-proven without real users and sustained traffic.
  • Results apply to the documented dataset, model, resolution, clip length, prompts, and compute envelope.

About

Compute-optimal data selection, adaptation, evaluation, and systems engineering for video-generation models.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages