Nebius
Now playingAfter the acquisition, the job became moving Tavily's production AI workloads onto Nebius AI Cloud without anyone on the outside noticing. Latency-sensitive services, so every requirement gets load tested before it reaches production.
Objectives
- Define Kubernetes networking, storage, observability, and reliability requirements for latency-sensitive services
- Build load testing and capacity planning workflows that find CPU, memory, networking, and scheduling bottlenecks before production does
- Standardize provisioning, releases, scaling, and incident response across AWS and Nebius
- Migrate production workloads to Nebius AI Cloud (in progress)
Full quest log
- Leading the post-acquisition migration of Tavily's production AI workloads to Nebius AI Cloud, defining Kubernetes networking, storage, observability, and reliability requirements for latency-sensitive services
- Built load testing and capacity planning workflows to validate production traffic targets and identify CPU, memory, networking, and scheduling bottlenecks before workloads reach production
- Standardized deployment and operational tooling across AWS and Nebius, making infrastructure provisioning, releases, scaling, and incident response repeatable across compute environments
- Partnered with Nebius compute and platform engineers to diagnose cross-layer performance and reliability issues spanning Kubernetes, networking, distributed systems, and application workloads