Parakh Jaggi

Senior Infrastructure Engineer

New York City

Press start

About

Hi, I'm Parakh. I'm an infrastructure engineer in New York, which means that when everything works nobody knows I exist, and when it doesn't I get very popular very fast. I've kept that act going across 150+ EKS clusters at FanDuel and the platform behind Tavily's AI search API, which I scaled from Series A through our acquisition by Nebius. These days I'm moving those production workloads onto Nebius AI Cloud, one load test at a time.

I like the parts of the stack most people would rather not think about: capacity planning, load testing, observability, cloud cost, and whatever paged at 3 am. The best infrastructure is the kind nobody notices, and I would rather delete a service than add one. If you want to compare notes on any of that, my inbox is open.

Work

LV 4

Nebius

Now playing

Senior Infrastructure Engineer, Team Lead New York City Feb 2026 – Present

After the acquisition, the job became moving Tavily's production AI workloads onto Nebius AI Cloud without anyone on the outside noticing. Latency-sensitive services, so every requirement gets load tested before it reaches production.

Objectives

  • Define Kubernetes networking, storage, observability, and reliability requirements for latency-sensitive services
  • Build load testing and capacity planning workflows that find CPU, memory, networking, and scheduling bottlenecks before production does
  • Standardize provisioning, releases, scaling, and incident response across AWS and Nebius
  • Migrate production workloads to Nebius AI Cloud (in progress)
Full quest log
  • Leading the post-acquisition migration of Tavily's production AI workloads to Nebius AI Cloud, defining Kubernetes networking, storage, observability, and reliability requirements for latency-sensitive services
  • Built load testing and capacity planning workflows to validate production traffic targets and identify CPU, memory, networking, and scheduling bottlenecks before workloads reach production
  • Standardized deployment and operational tooling across AWS and Nebius, making infrastructure provisioning, releases, scaling, and incident response repeatable across compute environments
  • Partnered with Nebius compute and platform engineers to diagnose cross-layer performance and reliability issues spanning Kubernetes, networking, distributed systems, and application workloads
LV 3

Tavily

Senior Infrastructure Engineer, Team Lead New York City Sept 2025 – Feb 2026 Acquired by Nebius

Joined as the infrastructure lead right after the $25M Series A and ran the platform through the Nebius acquisition. Latency-sensitive AI search on EKS, a team of three, and the pager for anything that mattered.

High scores

  1. 30K+ requests per minute sustained on core APIs, inside P99 latency SLOs
  2. 99.999% availability across production systems
  3. $200K+ cut from annual cloud spend through rightsizing and tuning
  4. 3 infrastructure engineers led
Full quest log
  • Scaled Tavily's infrastructure from its $25M Series A stage through acquisition by Nebius, evolving the platform to support rapid product and customer growth with zero infrastructure-related launch blockers
  • Designed and scaled a Kubernetes (EKS) platform for latency-sensitive AI search workloads, sustaining 30k+ requests per minute on core APIs while meeting P99 latency SLOs
  • Owned capacity planning and compute efficiency across production workloads, reducing annual cloud spend by $200k+ through rightsizing, resource tuning, and workload optimization without sacrificing performance
  • Built observability, alerting, and incident response practices across production systems, sustaining 99.999% availability and serving as the escalation point for high-severity infrastructure incidents
  • Operated and scaled distributed data systems including Redis, MongoDB, Elasticsearch, and Snowflake handling millions of daily queries with sub-second response times
  • Led a team of 3 infrastructure engineers, driving architecture reviews, platform standards, reliability improvements, and developer infrastructure initiatives
LV 2

FanDuel

Senior Platform Engineer New York City Sept 2024 – Sept 2025

Platform engineering at fleet scale: 150+ EKS clusters, their lifecycle and upgrades, and Compute API, the Go service that became the source of truth for all of them.

High scores

  1. 150+ EKS clusters under management
  2. 90% less onboarding time with standard Terraform modules and Helm charts
  3. 34% compute cost cut with Karpenter and custom scheduling
Full quest log
  • Led platform initiatives across 150+ Amazon EKS clusters, including cluster lifecycle management, migrations, upgrades, and onboarding for internal teams and applications
  • Led development of Compute API, a Go-based SQL-backed platform that serves as the source of truth for compute and Kubernetes cluster metadata across FanDuel
  • Designed and implemented Terraform modules and Helm charts to standardize and automate deployment of cluster utilities, improving consistency and reducing onboarding time by 90%
  • Built scalable CI/CD pipelines and GitOps workflows for cluster configuration management using ArgoCD, Terraform Cloud, and Helm
  • Optimized node provisioning and resource efficiency using Karpenter and custom scheduling strategies, reducing compute cost by 34%
  • Designed and implemented RESTful APIs using Huma, with data persistence via GORM and MySQL, enabling internal teams to query, manage, and audit cluster state programmatically
LV 1

Capital One

Senior Software Engineer New York City Oct 2021 – Sept 2024

Led a team of three building Empath, the app Capital One's customer service agents work in all day. Equal parts VueJS, AWS, and making releases boring.

High scores

  1. 25K agents at peak load
  2. 2M customer accounts serviced per week
  3. 50% faster releases with CodePipeline and CodeDeploy
  4. 20% faster page loads
Full quest log
  • Led a team of 3 engineers in the development of the microservice application "Empath"
  • Improved application efficiency by 30% through architecture optimization using VueJS, JavaScript, and AWS ECS
  • Implemented an AWS infrastructure that supported a peak user load of 25k agents, ensuring optimal scalability
  • Reduced release time by 50% by implementing CodePipeline and CodeDeploy for automated testing and deployment
  • Implemented VueJS performance improvements that reduced average page load time by 20%
  • Contributed to the launch of Empath, which enabled the successful servicing of 2 million customer accounts per week

Skills

parakh@nyc: ~ (it works, type something)
parakh@nyc:~$ kubectl get skills -A
NAMESPACE      NAME                STATUS    RESTARTS
kubernetes     eks                 Running   0
kubernetes     karpenter           Running   0
kubernetes     helm                Running   0
kubernetes     argocd              Running   0
kubernetes     istio               Running   0
kubernetes     cilium              Running   0
kubernetes     gpu-operator        Running   0
cloud          aws                 Running   0
cloud          nebius              Running   0
cloud          vpc                 Running   0
cloud          tailscale           Running   0
iac            terraform           Running   0
iac            github-actions      Running   0
iac            jenkins             Running   7
observability  cloudwatch          Running   0
observability  datadog             Running   0
observability  prometheus          Running   0
observability  opentelemetry       Running   0
reliability    load-testing        Running   0
reliability    capacity-planning   Running   0
reliability    on-call             Running   42
data           redis               Running   0
data           mongodb             Running   0
data           elasticsearch       Running   0
data           postgresql          Running   0
data           snowflake           Running   0
lang           go                  Running   0
lang           python              Running   0
lang           bash                Running   0
parakh@nyc:~$ 

Achievements

AWS Solutions Architect Professional

Unlocked Jan 2022

Baylor University

BS in Computer Science Waco, Texas. 2016 to 2020.

Contact