Description
We are seeking an exceptional Principal AI/ML Engineer to spearhead the development and optimization of our advanced machine learning models. In this role, you will design, implement, and fine-tune Large Language Models (LLMs) and diffusion models for a variety of complex tasks across natural language processing, computer vision, and speech processing. You will be responsible for building robust, end-to-end machine learning pipelines, from research and development to production deployment. This position requires deep expertise in transformer architectures, ASR systems, and modern MLOps practices. Collaborating with cross-functional teams, you will drive innovation, ensure the scalability and reliability of our AI solutions, and mentor team members, significantly influencing the technical direction of our projects.
Requirements
1. Advanced degree in Computer Science, Machine Learning, Data Science, or a related field.
2. 4+ years of professional experience in machine learning engineering with a strong focus on deep learning.
3. Expert-level proficiency in Python and extensive experience with deep learning frameworks such as PyTorch, TensorFlow, and the HuggingFace ecosystem.
4. Demonstrated expertise in designing, implementing, or fine-tuning Large Language Models (e.g., GPT, BERT, LLaMA).
5. Hands-on experience developing and optimizing diffusion models (e.g., Stable Diffusion, DDPM) or other generative models (VAEs, GANs).
6. Proven experience with Automatic Speech Recognition (ASR) systems like Whisper or Wav2Vec2 and speech processing pipelines.
7. Proficiency with parameter-efficient fine-tuning (PEFT) methods such as LoRA and QLoRA.
8. Solid understanding of MLOps practices and experience with tools like Docker, Kubernetes, MLflow, and cloud ML platforms.
Desirable
1. Experience with multimodal learning, combining text, audio, and visual data.
2. Knowledge of Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI.
3. Contributions to open-source ML projects or published research in top-tier conferences.
4. Experience with real-time inference systems and optimizing models for low-latency environments using techniques like quantization or TensorRT.
5. Familiarity with federated learning, privacy-preserving machine learning, and model interpretability techniques.