Description
We are seeking a visionary Principal AI/ML Engineer to lead the development and optimization of advanced machine learning models in natural language processing, computer vision, and speech processing. In this role, you will architect, implement, and fine-tune Large Language Models (LLMs) and diffusion models, shaping the future of our comprehensive data analytics platform. You will be responsible for building end-to-end ML pipelines, implementing state-of-the-art transformer architectures, and deploying scalable models into production. This position requires a strong research mindset to integrate cutting-edge techniques and a collaborative spirit to mentor team members. The ideal candidate is a driven expert in deep learning who thrives on solving complex challenges and is passionate about making a global impact through innovative AI solutions.
Requirements
1. Advanced degree (Master's or PhD) in Computer Science, Machine Learning, or a related field.
2. 4+ years of professional experience in machine learning engineering with a deep focus on deep learning and neural networks.
3. Expert-level Python skills and extensive experience with PyTorch, TensorFlow, and the HuggingFace ecosystem.
4. Demonstrated expertise in designing, implementing, and fine-tuning Large Language Models (e.g., GPT, BERT, LLaMA).
5. Hands-on experience with diffusion models (e.g., Stable Diffusion, DDPM) for generative tasks.
6. Proven experience with ASR systems (e.g., Whisper, Wav2Vec2) and building speech processing pipelines.
7. Solid understanding of parameter-efficient fine-tuning (PEFT) methods such as LoRA and QLoRA.
8. Familiarity with MLOps practices and tools (Docker, Kubernetes, MLflow) for model deployment and monitoring.
Desirable
1. Experience with multimodal learning, combining text, audio, and visual data.
2. Knowledge of Reinforcement Learning from Human Feedback (RLHF).
3. Experience developing real-time inference systems for streaming data.
4. Familiarity with model optimization techniques like quantization, pruning, or TensorRT.
5. Contributions to open-source ML projects or publications in top-tier AI conferences.