Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Agentic Mid-training Simulation Lab

中文说明

An engineering-oriented mini-lab for building and evaluating the data flywheel behind agentic foundation models. The project avoids full-scale foundation model training by combining API teacher models, synthetic trajectories, quality filters, curriculum sampling, lightweight LoRA/QLoRA recipes, and automatic evaluation.

The target use case is resume-ready research engineering work for roles focused on:

  • high-quality data systems and synthetic data
  • agentic planning and tool-use trajectories
  • long-context memory management
  • curriculum learning and data mixture analysis
  • lightweight remote GPU fine-tuning experiments
  • test-time scaling and verifier-based evaluation

What This Repo Demonstrates

seed tasks / long documents
  -> synthetic agentic trajectories
  -> quality filters and diversity checks
  -> curriculum sampler
  -> optional remote LoRA/QLoRA SFT
  -> tool-use, planning, and long-context evaluation
  -> experiment reports and resume bullets

The default path runs without network access by using a deterministic mock teacher. For real data generation, point the OpenAI-compatible provider at any hosted LLM API.

Quickstart

cd agentic-midtrain-lab
python -m venv .venv
source .venv/bin/activate
pip install -e .

agentic-lab generate --config configs/pipeline.json
agentic-lab filter --config configs/pipeline.json
agentic-lab sample --config configs/pipeline.json
agentic-lab eval --config configs/pipeline.json
agentic-lab report --config configs/pipeline.json

The commands write artifacts under outputs/:

  • outputs/synthetic/trajectories.jsonl
  • outputs/filtered/trajectories.filtered.jsonl
  • outputs/train/curriculum_train.jsonl
  • outputs/eval/eval_summary.json
  • outputs/reports/experiment_report.md

Optional API Teacher

Set these environment variables to use an OpenAI-compatible endpoint instead of the mock teacher:

export AGENTIC_LAB_PROVIDER=openai_compatible
export AGENTIC_LAB_API_KEY=...
export AGENTIC_LAB_BASE_URL=https://api.openai.com/v1
export AGENTIC_LAB_MODEL=gpt-4.1-mini

The generator stores all model outputs as JSONL so that the data can be inspected, filtered, versioned, and reused for training.

Remote GPU Fine-tuning

This repo does not require local model hosting. Use the generated training data on a remote GPU server:

rsync -av outputs/train/curriculum_train.jsonl user@gpu-box:/workspace/agentic-midtrain-lab/data/
rsync -av training/ user@gpu-box:/workspace/agentic-midtrain-lab/training/

Then run one of the LoRA/QLoRA recipes in training/configs/ with your preferred stack, such as Axolotl, TRL, PEFT, or your internal training launcher.

Project Layout

configs/
  pipeline.json                  # local pipeline config
data/raw/
  tasks_seed.jsonl               # seed agent tasks
  long_context_docs.jsonl        # toy long-context documents
src/agentic_midtrain_lab/
  cli.py                         # command line entrypoint
  synthetic.py                   # teacher-driven trajectory generation
  quality.py                     # scoring, validation, diversity filtering
  curriculum.py                  # difficulty/context/tool curriculum sampler
  memory.py                      # context pruning strategies
  agent_runtime.py               # small tool-use execution harness
  evaluation.py                  # automatic benchmark runner
  reporting.py                   # markdown report generator
training/
  configs/                       # remote LoRA/QLoRA experiment recipes
tests/
  test_pipeline.py               # unit tests for the pipeline

Resume Bullets

  • Built an agentic mid-training simulation lab covering synthetic planning data, tool-use trajectories, quality filtering, curriculum sampling, lightweight SFT recipes, and automated evaluation.
  • Designed rule-based and LLM-judge-ready filters for trajectory validity, tool alignment, diversity, and difficulty balancing, producing a structured agentic instruction dataset for LoRA/QLoRA experiments.
  • Implemented a long-context memory benchmark comparing full context, sliding window, summary memory, retrieval memory, and hybrid compression by accuracy, latency proxy, and token budget.
  • Created remote-GPU-ready training configs for baseline agentic SFT, curriculum SFT, and long-context SFT, enabling controlled ablations without local model hosting.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages