Principal ML Platform Engineer
Design and build scalable, reliable systems for training, serving, and operating generative AI models in production. Develop internal tooling and automation to reduce operational overhead for researchers and engineers, with a focus on agentic workflows and platform-level abstractions. Collaborate closely with research and product teams to improve observability, debugging, and developer experience across distributed GPU environments.