Compile ML research papers into optimized Triton kernels for agentskills.io.
- Constraint Extraction: Parses PDFs/markdown to extract math formulations and tensor requirements.
- AST Safety Sandbox: Validates generated code via AST inspection to block unsafe operations (
os,subprocess, etc.). - V2 Multi-File Synthesis: Separates kernels from PyTorch modules (
triton_kernel.pyvsnn_module.py). - Hardware-Aware Routing: Targets specific GPU architectures (
sm_80,sm_86,sm_90). - Hermes Integration: Generates compliant
SKILL.mdwith YAML frontmatter for automatic indexing.
By default, code execution is restricted. To enable local bare-metal execution (e.g., on cloud GPU instances):
export ALLOW_DANGER_RUN_BARE_METAL=trueWarning
This flag bypasses host isolation. Always verify generated code before execution.
pip install -e .Create a .env in the project root:
OPENAI_API_BASE="http://localhost:8000/v1"
OPENAI_API_KEY="your-api-key"
python -m paper_to_skill compile --pdf paper.txt --target sm_86 --out ./skills/python -m paper_to_skill install --dir ./skills/sageattention-2/Once installed, invoke the skill directly in Hermes:
hermes chat -q "/sageattention-2 'Run attention forward pass'"| Feature | Baseline LLM | Paper-to-Skill |
|---|---|---|
| Math Accuracy | Hallucinated/Placeholders | Extracted from Source |
| Quantization | Generic Fallback | INT8/FP8 per paper |
| Performance | Suboptimal/Non-runnable | Optimized Triton Kernel |