ML Platform & LLM Infrastructure Engineer
I build the infrastructure AI services run on — and measure whether it actually works.
Reliability is a number, not an adjective. Everything below was measured on a system I actually run.
| RAG answers that cite a source, or refuse | 80.6% → 100% |
| Strict retrieval accuracy · Hit@1 | 60.0% → 82.9% |
| Silent-failure detection time | 5.7 days → 40 min |
| Accelerator task execution time | ~70% faster than the default K8s scheduler |
| LLM serving under load | p95 18s single-request · throughput saturates at 4.3 req/min — bottleneck isolated to non-batched generation |
- Operating a multi-server GPU compute fleet (dual NVIDIA A40) serving LLM inference at the Korea Environment Institute
- Building KEIwi — an on-prem GPU fleet observability & incident platform
- Learning my way around Kubeflow, MLflow, Triton Inference Server, and feature stores
- Open to ML Platform / LLM Infrastructure roles
| Serving & inference | vLLM · Ollama · FastAPI · SSE streaming · GGUF quantization |
| Retrieval | Chroma · KURE-v1 embeddings · cross-encoder reranking |
| Orchestration | Kubernetes · Docker · containerd · systemd |
| Observability | Prometheus · Grafana · Loki · DCGM · OpenSearch |
| Cloud & edge | GCP · AWS · Oracle Cloud · Cloudflare Zero Trust |
| Languages | Python · TypeScript · C++ · C · Bash · SQL |
| KEIAdminSuperv | On-prem RAG over 363 internal regulation documents — every answer cites its source or refuses | FastAPI Ollama Chroma |
| KEIwi | GPU fleet observability & incident platform — fails safe: unknown state reports no-data, never a false down |
Prometheus Grafana DCGM |
| tpuserv | Accelerator-aware pod scheduling on bare-metal Kubernetes — published at KIISE 2024 | Kubernetes Coral TPU |
| dev-booth | Three LLM agents coordinating through a Kanban board only — a human keeps final merge authority | vLLM SQLite FastAPI |
| MineSweeper | Conflict-of-interest detection over mixed-format documents — it drafts, it never decides | Qwen2.5-VL Next.js |
Architecture diagrams and measured results for each → seanchoi.excusa.uk
Scheduling Techniques for Improving the Efficiency of TPU Work in a Cluster Environment — KIISE, 2024 · first author Research on Improving Pothole Object Detection Accuracy Through Similar-Object Data — KIISE, 2024 Proposal of a Korean History Education Application Based on Kubernetes Clustering — 2023 A Study on Chatbot for a Safe Harbor — 2023
Every project above was built alongside Claude — this is the receipt, not just the claim.






