Senior System Software Engineer - Local AI Automation

NVIDIA·Pune, Maharashtra, India·posted 4h ago · last seen 52m ago

Track this application

Get Started Free

Match score against your CV

Get Started Free

Tailor your resume to this job

Get Started Free

About interviewing at NVIDIA

Recruiter screen, a technical screen mixing resume deep-dive with live coding, a hiring-manager conversation, then a panel of three to five 45–60 minute rounds: coding, systems design under hardware constraints, a domain deep-dive, and behavioral. Highly team-specific — you interview directly with the team — with C++ depth expected almost universally and decisions sometimes taking five-plus weeks after the panel.

Read the full NVIDIA interview process →

Description

The Local AI Automation team is seeking a Senior System Software Engineer passionate about building and maintaining robust system-level software and infrastructure for deploying, benchmarking, and qualifying high-quality AI applications and models. You will collaborate with a team of highly qualified engineers dedicated to creating top-tier infrastructure for Local AI automation. If you have a real passion for building reliable, scalable automation systems in the AI ecosystem, this could be an excellent opportunity for us to work together!

What you'll be doing:

  • Design, implement, and maintain robust system software and infrastructure to efficiently run various AI workloads through different inference backends such as Llama.cpp, Ollama, PyTorch, vLLM, WinML, TRT-RTX, and others.
  • Scope out system requirements for deploying various AI applications, benchmarks, and models in automation, and develop solutions to measure correctness, functionality, reliability, and overall system performance.
  • Develop infrastructure to download AI models from various sources, efficiently manage a local model repository, and create automated synchronization mechanisms to keep the repository up to date.
  • Analyze large datasets and system workload metrics, implement data processing infrastructure to transform raw data into actionable information for developers, and design useful visualizations.
  • Collaborate with internal Local AI developers to identify and implement system-level instrumentation, diagnostics, and debugging capabilities that help quickly debug and isolate issues.
  • Design, implement, and maintain robust infrastructure for automated build, integration, deployment, and qualification across different inference frameworks and environments.
  • Understand existing infrastructure and system software architecture, debug issues in current frameworks, identify performance bottlenecks, and devise solutions to enhance system efficiency and scalability.

What we need to see:

  • 5+ years of experience with a B.Tech or higher degree in Computer Science, Information Technology, Software Engineering, or a related field.
  • Strong analytical, systems debugging, and problem-solving abilities, with the capacity to multitask effectively in a dynamic environment.
  • 5+ years of software development experience with strong programming skills in C/C++, C#, Java, or another systems programming language, along with exposure to at least one scripting language such as Python.
  • Experience with system APIs, multithreaded software, process and resource management, debugging, or performance analysis.
  • Familiarity with databases and SQL, source control systems, and CI/CD infrastructure such as Git, Perforce, and Jenkins.
  • Outstanding written and oral communication skills, enabling effective collaboration with management and engineering teams.

Ways to stand out from the crowd:

  • Experience building robust system software, runtime infrastructure, or scalable backend components.
  • Experience with system profiling, performance optimization, debugging, and telemetry/observability infrastructure, including analysis of Nsight Systems traces to understand system state and kernel performance.
  • Hands-on background with Linux/Windows system programming, Git CI/CD, databases/SQL, Kubernetes, Docker, and observability tools such as Grafana or Kibana.
  • Hands-on experience integrating or debugging inference frameworks such as Llama.cpp, Ollama, PyTorch, vLLM, WinML, or TRT-RTX.
  • Experience understanding CPU/GPU resource utilization, memory management, concurrency, and performance characteristics of AI workloads.

We are an equal-opportunity employer and value diversity at our company. With competitive salaries and a generous benefits package, we are widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us and, due to unprecedented growth, our exclusive engineering teams are rapidly growing. If you're a creative and autonomous engineer with a real passion for technology, we want to hear from you.

More engineering roles at NVIDIA