Use cases, pain points, and background
The all-in-one gym eval run flow currently rejects --input and requires a configured dataset split. To run an already-prepared JSONL file, users have to start servers separately with gym env start, then collect rollouts with gym eval run --no-serve --input.
Supporting explicit inputs in the all-in-one flow would simplify running subsets or retrying selected tasks without adding a separate dataset configuration.
Description:
Allow gym eval run --input <prepared-tasks.jsonl> to handle server startup, rollout collection, and shutdown in one command. Any required dataset preparation still happens beforehand.
Design:
- Update
nemo_gym/rollout_collection.py to accept an explicit input path without requiring a split.
- Update
nemo_gym/cli/eval.py to validate the file before server startup and use it directly, bypassing preparation of configured dataset splits.
- Give explicit input files precedence over configured splits, including when using a custom collection driver.
- Preserve existing split preparation and cache reuse when no input file is supplied.
- Update CLI help in
nemo_gym/cli/main.py, the CLI reference under fern/, and the relevant unit tests.
Out of scope:
- Automatically downloading or preparing the dataset supplied through
--input.
- Changing dataset formats, agent behavior, or verification logic.
- Changing the existing
--no-serve workflow.
Acceptance Criteria:
Use cases, pain points, and background
The all-in-one
gym eval runflow currently rejects--inputand requires a configured dataset split. To run an already-prepared JSONL file, users have to start servers separately withgym env start, then collect rollouts withgym eval run --no-serve --input.Supporting explicit inputs in the all-in-one flow would simplify running subsets or retrying selected tasks without adding a separate dataset configuration.
Description:
Allow
gym eval run --input <prepared-tasks.jsonl>to handle server startup, rollout collection, and shutdown in one command. Any required dataset preparation still happens beforehand.Design:
nemo_gym/rollout_collection.pyto accept an explicit input path without requiring a split.nemo_gym/cli/eval.pyto validate the file before server startup and use it directly, bypassing preparation of configured dataset splits.nemo_gym/cli/main.py, the CLI reference underfern/, and the relevant unit tests.Out of scope:
--input.--no-serveworkflow.Acceptance Criteria:
--splitor--no-serve.