About 1001
1001 builds AI-powered operational intelligence for complex, data-heavy environments. We turn fragmented data into a live, unified model of operations, helping governments and large enterprises make better decisions and solve high-stakes problems.
Our engagements begin with forward-deployed teams working inside customer environments. These teams use real operational data, build quickly, and iterate until the system proves itself. We then scale the system across the organization.
1001 is backed by Lux Capital, General Catalyst, CIV, Hanabi, Sanabil, and 9Yards, along with angels including Chris Ré, Amjad Masad, Karim Atiyeh, Kareem Amin, and Russell Kaplan.
About the role
Every model 1001 ships depends on the platform you build. As an ML Infrastructure Engineer, you will own the serving, deployment, and observability layer supporting models across multiple enterprise and government deployments.
You will make trained models production-ready by building infrastructure that is reliable, secure, and affordable to operate. The work includes multi-tenant serving, model registries and artifacts, deployment security, and efficient inference across CPU and GPU workloads.
This is a hands-on role for someone who has operated machine learning systems in production. You will also build shared platform components that product teams can use independently, making each deployment faster and more reliable than the last.
What you'll work on
- Build and operate ML serving infrastructure across multiple enterprise and government deployments.
- Own the deployment pipeline, including model registries, artifacts, datasets, deployment security, and multi-tenant serving.
- Build monitoring, logging, and observability systems that make model performance visible and surface problems early.
- Improve inference latency, throughput, reliability, and cost across CPU and GPU workloads.
- Maintain reusable platform components that allow product teams to deploy models without waiting for infrastructure support.
- Establish consistent deployment practices across customer environments.
What we're looking for
- At least 4 years of experience in infrastructure or platform engineering, including meaningful experience with ML infrastructure or MLOps.
- A track record of operating machine learning systems in production, not only developing models.
- Hands-on experience with a model serving stack such as Triton, TorchServe, Ray Serve, vLLM, BentoML, or KServe.
- Strong experience with Kubernetes, containers, CI/CD, and infrastructure as code using Terraform.
- Production cloud experience with AWS, GCP, or Azure.
- Experience with monitoring and observability tools such as Prometheus, Grafana, or OpenTelemetry.
- Strong Python and TypeScript skills.
Nice to have
- Experience optimizing GPU workloads.
- Experience with distributed training.
- Experience serving large models in production.
Working at 1001
We take on high-stakes problems in environments where mistakes carry real consequences. That demands an uncompromising bar, real speed, and systems that hold up under live operations. The people who thrive set that bar for themselves and keep raising it. They own outcomes end to end, bring rigor to everything, and lift everyone around them.