Skip to content
View mooner92's full-sized avatar
:shipit:
Bee's Knees
:shipit:
Bee's Knees

Block or report mooner92

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
mooner92/README.md

Sean Choi  ·  최명헌

ML Platform & LLM Infrastructure Engineer
I build the infrastructure AI services run on — and measure whether it actually works.

Portfolio  ·  LinkedIn  ·  Email


Reliability is a number, not an adjective. Everything below was measured on a system I actually run.

Measured — not claimed

RAG answers that cite a source, or refuse 80.6% → 100%
Strict retrieval accuracy · Hit@1 60.0% → 82.9%
Silent-failure detection time 5.7 days → 40 min
Accelerator task execution time ~70% faster than the default K8s scheduler
LLM serving under load p95 18s single-request · throughput saturates at 4.3 req/min — bottleneck isolated to non-batched generation

Currently

  • Operating a multi-server GPU compute fleet (dual NVIDIA A40) serving LLM inference at the Korea Environment Institute
  • Building KEIwi — an on-prem GPU fleet observability & incident platform
  • Learning my way around Kubeflow, MLflow, Triton Inference Server, and feature stores
  • Open to ML Platform / LLM Infrastructure roles

Stack

Serving & inference vLLM · Ollama · FastAPI · SSE streaming · GGUF quantization
Retrieval Chroma · KURE-v1 embeddings · cross-encoder reranking
Orchestration Kubernetes · Docker · containerd · systemd
Observability Prometheus · Grafana · Loki · DCGM · OpenSearch
Cloud & edge GCP · AWS · Oracle Cloud · Cloudflare Zero Trust
Languages Python · TypeScript · C++ · C · Bash · SQL

Projects

KEIAdminSuperv On-prem RAG over 363 internal regulation documents — every answer cites its source or refuses FastAPI Ollama Chroma
KEIwi GPU fleet observability & incident platform — fails safe: unknown state reports no-data, never a false down Prometheus Grafana DCGM
tpuserv Accelerator-aware pod scheduling on bare-metal Kubernetes — published at KIISE 2024 Kubernetes Coral TPU
dev-booth Three LLM agents coordinating through a Kanban board only — a human keeps final merge authority vLLM SQLite FastAPI
MineSweeper Conflict-of-interest detection over mixed-format documents — it drafts, it never decides Qwen2.5-VL Next.js

Architecture diagrams and measured results for each → seanchoi.excusa.uk


Publications

Scheduling Techniques for Improving the Efficiency of TPU Work in a Cluster Environment — KIISE, 2024 · first author Research on Improving Pothole Object Detection Accuracy Through Similar-Object Data — KIISE, 2024 Proposal of a Korean History Education Application Based on Kubernetes Clustering — 2023 A Study on Chatbot for a Safe Harbor — 2023


Every project above was built alongside Claude — this is the receipt, not just the claim.

Tokscale Stats

Pinned Loading

  1. tpuserv tpuserv Public

    클러스터 환경에서의 TPU 작업 효율 향상 스케줄링 기법 연구

    JavaScript 1

  2. dev-booth dev-booth Public

    dev-booth — Autonomous Multi-Agent Development System

    Python 1

  3. KEIAdminSuperv KEIAdminSuperv Public

    Horong — On-Prem RAG Chatbot for Internal Regulations

    Python 1

  4. KEIwi KEIwi Public

    KEIwi — GPU Fleet Observability & Incident Platform

    Python

  5. MineSweeper MineSweeper Public

    MineSweeper — Conflict-of-Interest Detection for Recruitment

    TypeScript

  6. Raypoke Raypoke Public

    Advanced Rayban glasses by using poke

    TypeScript