GPU Kernel Engineer (OpenAI Triton / CUDA C++)
Deevo Systems & Co.
Remote
Role Description The GPU Kernel Engineer (OpenAI Triton / CUDA C++) will design, implement, and optimize high‑performance GPU kernels that power Deevo’s foundation models, context horizon systems, and autonomous multi‑agent platforms. Day‑to‑day work includes writing and tuning Triton and CUDA C++ kernels, profiling and optimizing memory access patterns and FLOP utilization, and integrating low‑level kernels into inference and training pipelines. The role involves collaborating with ML researchers and systems engineers to translate model and systems requirements into efficient GPU implementations, debugging performance issues across hardware and software layers, and contributing to internal tooling for benchmarking and kernel deployment. This is a full‑time remote role, with asynchronous collaboration across time zones and a strong focus on disciplined, production‑grade engineering.
Qualifications
- Strong proficiency in GPU programming and performance engineering, including OpenAI Triton, CUDA C++, and experience optimizing kernels for latency, throughput, and memory efficiency.
- Solid systems and low‑level engineering skills, such as C/C++ development, understanding of GPU architectures (SMs, warps, memory hierarchy), and experience with profiling tools (e.g., Nsight, nvprof, perf).
- Background in numerical computing and machine learning systems, including matrix/tensor operations, attention mechanisms, and experience integrating kernels into deep learning frameworks or custom runtimes.
- Experience with high‑performance and distributed compute environments, such as multi‑GPU or cluster‑based training/inference, and familiarity with concepts like ring‑based communication and memory routing.
- Ability to write clean, tested, and maintainable code, including version control (Git), code review practices, and documentation for low‑level components and performance benchmarks.
- Strong analytical and problem‑solving skills, with the ability to diagnose bottlenecks across the stack