company-logo
Principal Applied AI Architect
Description
A client of bytespark.ai is seeking a Principal Applied AI Architect to lead the strategy, architecture, and delivery of advanced AI solutions for high-volume voice, text, and multimodal data processing. The successful candidate will own end-to-end technical decisions for large language models, speech recognition, diffusion models, and production machine learning platforms. This role will design and fine-tune transformer-based systems for classification, entity extraction, sentiment analysis, conversational AI, and other domain-specific applications. The architect will establish scalable approaches for distributed training, model and hardware sizing, parameter-efficient fine-tuning, inference optimization, and real-time deployment. They will guide the development of robust pipelines that process billions of data points while meeting demanding reliability, security, governance, and performance expectations. The position will partner with engineering, product, research, and operational stakeholders to translate complex requirements into maintainable AI architectures and delivery roadmaps. Responsibilities include defining evaluation frameworks, monitoring model quality, directing A/B testing, and promoting reproducible, explainable, and ethical AI practices. The architect will document standards, mentor machine learning engineers, review critical designs, and advance engineering practices across multiple teams. This is an opportunity for a hands-on technical leader to shape consequential AI systems while continuously evaluating emerging research and technologies.
Requirements
1. Advanced degree in Computer Science, Machine Learning, Data Science, Mathematics, or a related discipline, with at least 4 years of machine learning engineering experience focused on deep learning and neural networks.
2. Principal-level experience defining end-to-end applied AI architecture, technical strategy, governance standards, design patterns, and delivery roadmaps for high-impact production programs.
3. Expert Python proficiency and advanced hands-on experience with PyTorch, TensorFlow or JAX, HuggingFace Transformers, and at least one performance-oriented language such as C, C++, JavaScript, or Julia.
4. Demonstrated delivery of transformer and LLM solutions using architectures such as GPT, BERT, T5, or LLaMA, including fine-tuning for NLP and conversational AI use cases.
5. Production experience with ASR and speech-processing pipelines using technologies such as Whisper, Wav2Vec2, torchaudio, librosa, or SpeechBrain.
6. Experience sizing, optimizing, and deploying large models using distributed training, model parallelism, quantization, pruning, distillation, ONNX, TensorRT, or comparable methods.
7. Ability to build distributed multimodal pipelines and operate models with Docker, Kubernetes, MLflow, Weights & Biases, cloud ML platforms, Git, DVC, Spark, Dask, or Ray.
8. Strong technical leadership, English communication, documentation, stakeholder collaboration, design review, and mentoring capabilities.
Desirable
1. Hands-on experience with diffusion models, VAEs, GANs, synthetic data generation, and multimodal content generation.
2. Knowledge of RLHF, constitutional AI, federated learning, privacy-preserving ML, explainable AI, bias detection, or adversarial machine learning.
3. Experience building real-time inference and streaming architectures for telecommunications, public-sector, cybersecurity, or similarly regulated environments.
4. Published AI research, patents, conference participation, or meaningful contributions to open-source machine learning projects.
5. Experience with edge deployment, mobile model optimization, graph neural networks, knowledge graphs, AutoML, or neural architecture search.
Getting StartedA few quick details so we know how to reach you
How did you hear about us? *
Which country's passport do you hold? *
Email *(Please ensure the email matches the one mentioned in your CV or resume)
LinkedIn Profile URL *
Please mention your notice period *
Let’s Get to Know You BetterA few short questions to understand your experience and what you enjoy doing
1. Do you have an advanced degree in a relevant quantitative discipline and at least 4 years of machine learning engineering experience focused on deep learning? *
2. Have you owned the end-to-end architecture and technical strategy of at least one production applied AI platform at Principal, Lead, or equivalent level? *
3. Have you implemented and fine-tuned transformer-based LLMs using Python, PyTorch, TensorFlow or JAX, and HuggingFace Transformers? *
4. Have you deployed a production speech-to-text pipeline using Whisper, Wav2Vec2, or a comparable ASR model? *
5. Have you sized and optimized large models for production using distributed training, model parallelism, quantization, pruning, distillation, ONNX, TensorRT, or similar techniques? *
6. Have you operated production ML systems using container orchestration, experiment tracking, model monitoring, version control, and reproducible deployment workflows? *
Final DetailsSalary expectations and any supporting credentials
1. Where does your salary sit today (so we can help it move up tomorrow)?*
Enter your monthly salary in your local currency
2. What’s the number that’ll make you say "this is worth it"?*
Per month, in the currency mentioned
Upload ResumeHelp us get to know you better by sharing your most recent resume