Bridging Agentic AI
to System Reality

Agent-Native Inference OS —
AI, SW, and HW on Memory-centric Infrastructure.


Co-Founders

Dongsoo Lee
Chief Executive Officer
Dongsoo Lee|이동수
  • National AI Strategy Committee Member | 국가인공지능전략위원회 위원

  • Ph.D. @ Purdue University·Memory Design, VLSI Testing, Emerging Memory
  • Research Staff Member @ IBM TJ Watson Research·Power9/10 CPU, NPU Design, FP8 Training
  • Principal Engineer @ Samsung Research·AI Model Compression Lead
  • EVP @ NAVER Cloud·AI Computing Strategy and Efficient Serving Systems
Minsoo Rhu
Chief Research Officer
Minsoo Rhu|유민수
  • Endowed Chair Professor @ KAIST·AI Computing System / GPU & NPU HW/SW Design
  • National AI Strategy Committee Member | 국가인공지능전략위원회 위원

  • Ph.D. @ University of Texas, Austin·GPU Processor and Memory System Design
  • Research Scientist @ Meta·Hardware/Software Systems for Secure and Private AI
  • Senior Research Scientist @ NVIDIA·GPU Memory Virtualization / AI Accelerator Design
  • Program Chair for MICRO 2025, MICRO/ISCA/HPCA Hall of Fame
Baeseong Park
Chief Product Officer
Baeseong Park|박배성

  • Engineer @ Samsung Research·Optimized Transformer Kernel
  • Leader @ NAVER Cloud·AI Serving System
Se Jung Kwon
Chief Strategy Officer
Se Jung Kwon|권세중

  • Ph.D. @ KAIST
  • Staff Engineer @ Samsung Research·Model Compression
  • Leader @ NAVER Cloud·AI-HW Partnership / Strategy
  • Adjunct Professor @ KAIST·NAVER-INTEL-KAIST AI Research Center

Build With Us

We're not hiring for headcount — we're looking for people who want to take on the Agent-native challenge from the ground up.
These are the roles we're opening now:

a2sys는 빈자리를 채우려고 채용하지 않습니다. AI 시스템을 밑바닥부터 다시 설계하는 일에 함께 도전할 인재를 찾고 있습니다.
다음은 현재 채용 중인 포지션입니다.

Role What You'll Own
Agent Platform
Engineer
에이전트 플랫폼 엔지니어

You build the backend and runtime of the platform that defines and executes agents. In an agent product, the reliability of execution shapes quality as much as model performance does. You own the stability of the execution layer and the traceability of every run.

Agents behave differently on the same input, and the ways they fail are hard to predict. You design how execution state, memory, tool calls, and model calls are recorded, so that anyone can retrace what went wrong and where.

에이전트를 정의하고 실행하는 플랫폼의 백엔드와 런타임을 만드는 역할입니다. 에이전트 제품에서는 모델 성능만큼 실행의 신뢰성이 품질을 좌우합니다. 실행 계층의 안정성과 실행 추적을 책임집니다.

에이전트는 같은 입력에도 다르게 움직이고 실패하는 방식도 예측하기 어렵습니다. 실행 상태와 메모리, 도구 호출과 모델 호출의 기록을 설계해 무엇이 어디서 잘못됐는지 되짚을 수 있게 만듭니다.

LLM Modeling
Engineer
모델링 엔지니어

You sharpen a model's intelligence to meet product requirements and own the full model lifecycle, so it runs at a cost and speed we can actually serve. The biggest model is not always the answer. We aim for small, sharp, specialized models that clear the quality bar on as little GPU as possible.

Depending on your strengths and experience, you take the lead on one of two tracks — Agent Fine-tuning or Model Compression & Optimization — working closely with our infrastructure and inference engineers.

모델의 지능을 제품 요구사항에 맞게 고도화하고, 서비스 가능한 비용과 속도로 실행될 수 있도록 모델 라이프사이클 전반을 책임지는 역할입니다. 가장 거대한 모델이 항상 정답은 아닙니다. 우리 팀은 요구되는 품질을 달성하면서도 GPU 자원을 최소화할 수 있는 '똑똑하고 가벼운 특화 모델'을 지향합니다.

본 포지션은 지원자의 강점과 경험에 따라 [에이전트 파인튜닝] 또는 [모델 경량화 및 최적화] 중 하나의 핵심 역할을 수행하며, 인프라 및 추론 엔지니어와 긴밀하게 협업합니다.

LLM Inference
Engineer
추론 엔지니어

You run and optimize the serving stack yourself, built on inference engines such as vLLM, SGLang, and TensorRT-LLM. In an environment where inference performance and GPU utilization are the company's cost competitiveness, you own the serving layer's throughput, latency, and cost per token.

It does not stop at tuning engine settings. You trace the execution path down to the source to find bottlenecks, and turn what you measure into serving configurations and GPU capacity plans.

vLLM, SGLang, TensorRT-LLM 등 추론 엔진 기반의 서빙 스택을 직접 운영하고 최적화하는 역할입니다. 추론 성능과 GPU 활용률이 곧 회사의 원가 경쟁력이 되는 환경에서, 서빙 레이어의 처리량과 지연시간, 토큰당 비용을 책임집니다.

엔진 설정 튜닝에서 끝나지 않습니다. 실행 경로를 소스 수준까지 분석해 병목을 찾고, 측정 결과를 서빙 구성과 GPU 용량 계획으로 연결합니다.

AI Infrastructure
Engineer
AI 인프라 엔지니어

You keep AI services running without interruption on Kubernetes and GPU clusters. GPUs are the most expensive asset the company owns, and when the service stops, the product stops with it. You own the reliability and resource efficiency of our infrastructure.

Clusters and nodes, deployment automation, networking and security, observability — all of it is yours. Tuning inference engine performance belongs to the LLM Inference Engineer; this role builds the ground underneath, so every other engineer can put their workload on it safely.

Kubernetes와 GPU 클러스터 위에서 AI 서비스가 끊기지 않고 돌아가게 만드는 역할입니다. GPU는 회사에서 가장 비싼 자산이고, 서비스가 멈추면 제품도 멈춥니다. 인프라의 안정성과 자원 효율을 책임집니다.

클러스터와 노드, 배포 자동화, 네트워크와 보안, 관측 체계를 소유합니다. 추론 엔진의 성능 튜닝은 LLM Inference Engineer가 맡고, 이 역할은 그 위에서 다른 엔지니어가 자기 워크로드를 안전하게 올릴 수 있는 기반을 만듭니다.

Full job descriptions on our careers page. Don't see your role? We'd still love to hear from you.
자세한 업무 내용과 자격 요건은 채용 페이지에서 확인하실 수 있습니다.


Contact

Questions about a2sys, hiring, or partnerships? Send us a message.
회사, 채용, 파트너십 관련 문의는 편하게 남겨 주세요.

Hiring Partnership & Business Technical Collaboration General Inquiry
Get in Touch