Deep Learning Performance Software Intern - 2027

NVIDIA·Shanghai, Shanghai, China | Beijing, Beijing, China·posted 3d ago · last seen 52m ago

Track this application

Get Started Free

Match score against your CV

Get Started Free

Tailor your resume to this job

Get Started Free

About interviewing at NVIDIA

Recruiter screen, a technical screen mixing resume deep-dive with live coding, a hiring-manager conversation, then a panel of three to five 45–60 minute rounds: coding, systems design under hardware constraints, a domain deep-dive, and behavioral. Highly team-specific — you interview directly with the team — with C++ depth expected almost universally and decisions sometimes taking five-plus weeks after the panel.

Read the full NVIDIA interview process →

Description

We are now looking for a Deep Learning Performance Software Engineering Intern!
 
We are expanding our research and development for deep learning. We seek excellent Software Engineers to join our team. We specialize in developing GPU-accelerated Deep learning software. Researchers around the world are using NVIDIA GPUs to power a revolution in deep learning, enabling breakthroughs in numerous areas. Join the team that builds software to enable new solutions. Your ability to work in a fast-paced customer-oriented team is required and excellent communication skills are necessary. 
 
What you’ll be doing:

  • Creating and maintaining SKILL, Wiki, and agent harness

  • Develop TileGym, Triton CUDA TileIR backend and CUDA Tile

  • Develop highly optimized deep learning kernels through tile-based GPU programming model

  • End-to-end performance optimization through tile-based GPU programming model

  • Do performance optimization, analysis, and tuning

 
What we need to see: 

  • Pursuing a degree from a university in an engineering or computer science related field. A masters or doctoral candidate is preferred.

  • Understands the core components of agentic systems, including LLM APIs, prompting, tool use/function calling, agent loops, planning, reasoning, memory, RAG, skills, MCP, sub-agents, multi-agent architectures, context engineering, and harness engineering.

  • Excellent C/C++ programming and software design skills

  • Python experience a plus

  • MLIR experience a plus

  • Performance modelling, profiling, debug, and code optimization or architectural knowledge of CPU and GPU

  • GPU programming experience (CUDA or OpenCL) desired

 
NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most brilliant and talented people on the planet working for us. If you're creative and autonomous, we want to hear from you!

More engineering roles at NVIDIA