Senior System Software Engineer - Local AI Automation
Track this application
Get Started FreeMatch score against your CV
Get Started FreeTailor your resume to this job
Get Started FreeAbout interviewing at NVIDIA
Recruiter screen, a technical screen mixing resume deep-dive with live coding, a hiring-manager conversation, then a panel of three to five 45–60 minute rounds: coding, systems design under hardware constraints, a domain deep-dive, and behavioral. Highly team-specific — you interview directly with the team — with C++ depth expected almost universally and decisions sometimes taking five-plus weeks after the panel.
Read the full NVIDIA interview process →Description
The Local AI Automation team is seeking a Senior System Software Engineer passionate about building and maintaining robust system-level software and infrastructure for deploying, benchmarking, and qualifying high-quality AI applications and models. You will collaborate with a team of highly qualified engineers dedicated to creating top-tier infrastructure for Local AI automation. If you have a real passion for building reliable, scalable automation systems in the AI ecosystem, this could be an excellent opportunity for us to work together!
What you'll be doing:
- Design, implement, and maintain robust system software and infrastructure to efficiently run various AI workloads through different inference backends such as Llama.cpp, Ollama, PyTorch, vLLM, WinML, TRT-RTX, and others.
- Scope out system requirements for deploying various AI applications, benchmarks, and models in automation, and develop solutions to measure correctness, functionality, reliability, and overall system performance.
- Develop infrastructure to download AI models from various sources, efficiently manage a local model repository, and create automated synchronization mechanisms to keep the repository up to date.
- Analyze large datasets and system workload metrics, implement data processing infrastructure to transform raw data into actionable information for developers, and design useful visualizations.
- Collaborate with internal Local AI developers to identify and implement system-level instrumentation, diagnostics, and debugging capabilities that help quickly debug and isolate issues.
- Design, implement, and maintain robust infrastructure for automated build, integration, deployment, and qualification across different inference frameworks and environments.
- Understand existing infrastructure and system software architecture, debug issues in current frameworks, identify performance bottlenecks, and devise solutions to enhance system efficiency and scalability.
What we need to see:
- 5+ years of experience with a B.Tech or higher degree in Computer Science, Information Technology, Software Engineering, or a related field.
- Strong analytical, systems debugging, and problem-solving abilities, with the capacity to multitask effectively in a dynamic environment.
- 5+ years of software development experience with strong programming skills in C/C++, C#, Java, or another systems programming language, along with exposure to at least one scripting language such as Python.
- Experience with system APIs, multithreaded software, process and resource management, debugging, or performance analysis.
- Familiarity with databases and SQL, source control systems, and CI/CD infrastructure such as Git, Perforce, and Jenkins.
- Outstanding written and oral communication skills, enabling effective collaboration with management and engineering teams.
Ways to stand out from the crowd:
- Experience building robust system software, runtime infrastructure, or scalable backend components.
- Experience with system profiling, performance optimization, debugging, and telemetry/observability infrastructure, including analysis of Nsight Systems traces to understand system state and kernel performance.
- Hands-on background with Linux/Windows system programming, Git CI/CD, databases/SQL, Kubernetes, Docker, and observability tools such as Grafana or Kibana.
- Hands-on experience integrating or debugging inference frameworks such as Llama.cpp, Ollama, PyTorch, vLLM, WinML, or TRT-RTX.
- Experience understanding CPU/GPU resource utilization, memory management, concurrency, and performance characteristics of AI workloads.
We are an equal-opportunity employer and value diversity at our company. With competitive salaries and a generous benefits package, we are widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us and, due to unprecedented growth, our exclusive engineering teams are rapidly growing. If you're a creative and autonomous engineer with a real passion for technology, we want to hear from you.