- Updated: February 27, 2026
- 5 min read
Founding Reliability & Performance Engineer Role at LiteLLM – Join the AI Infrastructure Revolution
LiteLLM is hiring a Founding Reliability & Performance Engineer to safeguard its high‑traffic AI gateway, ensuring millions of daily LLM API calls run smoothly for customers like NASA, Adobe, and Netflix.
Founding Reliability & Performance Engineer at LiteLLM – Join the AI Infrastructure Frontier
LiteLLM, the open‑source AI gateway with over 36K GitHub stars, is expanding its core team. The company seeks a senior engineer who will own reliability, performance, and production stability for a platform that routes hundreds of millions of large‑language‑model (LLM) API calls each day. This role blends hands‑on incident response with deep performance engineering, offering a rare chance to shape the backbone of modern AI workloads.

Who Is LiteLLM?
Founded in 2023 and backed by Y Combinator’s W23 batch, LiteLLM provides a Python SDK and proxy server that normalizes more than 100 LLM providers into a single OpenAI‑compatible API. With $7 M ARR and a lean 10‑person team, the startup powers AI‑intensive workloads for enterprises such as NASA, Adobe, Netflix, Stripe, and Nvidia. Its rapid growth means any outage directly impacts mission‑critical AI pipelines, making reliability a top priority.
What You’ll Own: Reliability & Performance Responsibilities
The role is roughly 60 % operational reliability and 40 % performance engineering. Expect a dynamic mix of tasks that keep the platform healthy and fast.
- Lead on‑call rotations, triage critical incidents, and produce blameless post‑mortems.
- Design and maintain self‑healing mechanisms for
PostgreSQLandRedisoutages. - Develop soak tests and CI‑integrated load simulations to catch regressions before release.
- Profile and eliminate memory leaks in long‑running Python
asyncioservices, targetingOOMprevention after hours of sustained load. - Optimize hot‑path request handling to keep added latency under 10 ms at 5K+ RPS (P50/P95/P99 benchmarks).
- Implement structured logging, distributed tracing, and accurate Prometheus metrics for real‑time observability.
- Collaborate with product and engineering to define SLOs for enterprise customers and build automated canary deployments with rollback.
- Review PRs that affect the request pipeline, quantifying performance impact (e.g., “+50 ms at P99”) and recommending mitigations.
Required Qualifications & Technical Skills
LiteLLM looks for a candidate who thrives in high‑scale, low‑latency environments and enjoys solving “hard‑to‑debug” problems.
- 2+ years of production experience with Python services, especially
asyncioandaiohttp/httpx. - Deep understanding of Python async internals, connection pooling, and memory‑management tools (e.g.,
memray,py‑spy,tracemalloc). - Strong PostgreSQL expertise: connection‑pool tuning, query optimization, and handling high‑throughput write patterns.
- Hands‑on Kubernetes knowledge: pod lifecycle, resource limits, health probes, and troubleshooting at the cluster level.
- Proven on‑call experience with a calm, systematic approach to incident resolution.
- Experience with API gateways, proxies, or load balancers where overhead is a primary metric.
- Familiarity with HTTP/2, Server‑Sent Events (SSE), and streaming response handling in async Python.
- Bonus: Contributions to open‑source infrastructure projects or prior roles at companies like Meta, Cloudflare, Fastly, Datadog, or Stripe.
Compensation, Benefits, & Equity
LiteLLM offers a competitive total‑compensation package designed for senior engineers in the AI space.
| Component | Details |
|---|---|
| Base Salary | $200K – $270K per year |
| Equity | 0.25 % – 0.75 % (meaningful at early‑stage growth) |
| Health & Wellness | Comprehensive medical, dental, vision, and mental‑health benefits |
| Remote Flexibility | Fully remote or hybrid options in the San Francisco Bay Area |
| Professional Growth | Conference budget, continuous learning stipend, and mentorship from founders |
How to Apply – Leverage UBOS Resources
Applying is straightforward, but we recommend exploring UBOS tools that can streamline your workflow and showcase your expertise.
- Visit the UBOS homepage to learn about the platform that powers many AI‑centric SaaS products.
- Review the UBOS platform overview for insights into low‑code AI integration.
- Explore the Enterprise AI platform by UBOS to see how large‑scale AI workloads are managed.
- Check out the Workflow automation studio for building reliable pipelines—experience that aligns with LiteLLM’s needs.
- Use the Web app editor on UBOS to prototype monitoring dashboards you could later adapt for LiteLLM.
- Browse the UBOS templates for quick start and consider the AI SEO Analyzer template as a showcase of performance‑focused AI tooling.
- Read about the UBOS partner program to understand how strategic collaborations are structured.
- Finally, submit your resume and a brief cover letter through the official posting: LiteLLM Founding Reliability & Performance Engineer job.
Why This Position Accelerates Your Tech Career
Being the first dedicated reliability hire at a high‑growth AI startup offers unparalleled visibility:
- Impact at Scale: Your work directly affects millions of daily AI calls for marquee customers.
- Open‑Source Credibility: Contributions will be visible on a repository with 36K stars, boosting your personal brand.
- Ownership & Autonomy: No bureaucratic layers—identify a problem, design a solution, ship it.
- Equity Upside: Early‑stage equity can become a substantial financial reward as LiteLLM scales.
- Cross‑Domain Mastery: You’ll deepen expertise in Python async, Kubernetes, PostgreSQL, and AI‑specific networking.
Take the Next Step
If you thrive on solving complex performance puzzles, love the idea of keeping a critical AI gateway humming, and want to leave a lasting imprint on the AI infrastructure landscape, this is your moment. Apply now, and join a team that’s redefining how the world accesses large‑language models.
Andrii Bidochko
CTO UBOS
Andrii Bidochko is an AI entrepreneur and researcher focused on AI agents, reinforcement learning, and autonomous systems. He writes about the technologies shaping the future of machine intelligence, from frontier models and agent architectures to real-world AI applications.