An engineering-oriented mini-lab for building and evaluating the data flywheel behind agentic foundation models. The project avoids full-scale foundation model training by combining API teacher models, synthetic trajectories, quality filters, curriculum sampling, lightweight LoRA/QLoRA recipes, and automatic evaluation.
The target use case is resume-ready research engineering work for roles focused on:
- high-quality data systems and synthetic data
- agentic planning and tool-use trajectories
- long-context memory management
- curriculum learning and data mixture analysis
- lightweight remote GPU fine-tuning experiments
- test-time scaling and verifier-based evaluation
seed tasks / long documents
-> synthetic agentic trajectories
-> quality filters and diversity checks
-> curriculum sampler
-> optional remote LoRA/QLoRA SFT
-> tool-use, planning, and long-context evaluation
-> experiment reports and resume bullets
The default path runs without network access by using a deterministic mock teacher. For real data generation, point the OpenAI-compatible provider at any hosted LLM API.
cd agentic-midtrain-lab
python -m venv .venv
source .venv/bin/activate
pip install -e .
agentic-lab generate --config configs/pipeline.json
agentic-lab filter --config configs/pipeline.json
agentic-lab sample --config configs/pipeline.json
agentic-lab eval --config configs/pipeline.json
agentic-lab report --config configs/pipeline.jsonThe commands write artifacts under outputs/:
outputs/synthetic/trajectories.jsonloutputs/filtered/trajectories.filtered.jsonloutputs/train/curriculum_train.jsonloutputs/eval/eval_summary.jsonoutputs/reports/experiment_report.md
Set these environment variables to use an OpenAI-compatible endpoint instead of the mock teacher:
export AGENTIC_LAB_PROVIDER=openai_compatible
export AGENTIC_LAB_API_KEY=...
export AGENTIC_LAB_BASE_URL=https://api.openai.com/v1
export AGENTIC_LAB_MODEL=gpt-4.1-miniThe generator stores all model outputs as JSONL so that the data can be inspected, filtered, versioned, and reused for training.
This repo does not require local model hosting. Use the generated training data on a remote GPU server:
rsync -av outputs/train/curriculum_train.jsonl user@gpu-box:/workspace/agentic-midtrain-lab/data/
rsync -av training/ user@gpu-box:/workspace/agentic-midtrain-lab/training/Then run one of the LoRA/QLoRA recipes in training/configs/ with your preferred
stack, such as Axolotl, TRL, PEFT, or your internal training launcher.
configs/
pipeline.json # local pipeline config
data/raw/
tasks_seed.jsonl # seed agent tasks
long_context_docs.jsonl # toy long-context documents
src/agentic_midtrain_lab/
cli.py # command line entrypoint
synthetic.py # teacher-driven trajectory generation
quality.py # scoring, validation, diversity filtering
curriculum.py # difficulty/context/tool curriculum sampler
memory.py # context pruning strategies
agent_runtime.py # small tool-use execution harness
evaluation.py # automatic benchmark runner
reporting.py # markdown report generator
training/
configs/ # remote LoRA/QLoRA experiment recipes
tests/
test_pipeline.py # unit tests for the pipeline
- Built an agentic mid-training simulation lab covering synthetic planning data, tool-use trajectories, quality filtering, curriculum sampling, lightweight SFT recipes, and automated evaluation.
- Designed rule-based and LLM-judge-ready filters for trajectory validity, tool alignment, diversity, and difficulty balancing, producing a structured agentic instruction dataset for LoRA/QLoRA experiments.
- Implemented a long-context memory benchmark comparing full context, sliding window, summary memory, retrieval memory, and hybrid compression by accuracy, latency proxy, and token budget.
- Created remote-GPU-ready training configs for baseline agentic SFT, curriculum SFT, and long-context SFT, enabling controlled ablations without local model hosting.