التخصص
الهندسة
نوع الدوام
دوام كامل
الخبرة
أي مستوى
الموقع
أبوظبي
وصف الوظيفة
About AI71:AI71 is an industry leader in artificial intelligence, delivering innovative solutions that empower developers, businesses and governments to solve complex challenges. AI71 builds secure, enterprise-ready applications powered by cutting-edge technology—tailored for knowledge workers and sector-specific needs. AI71 bridges the gap between advanced AI and real-world impact. Guided by a strong commitment to research and responsibility, we create transformative solutions that drive progress and empower communities.The Role:As an MLOps Engineer you set the ML infrastructure and reliability strategy across AI71's platform, including how LLMs and other deep learning models are deployed, fine-tuned, and served at scale. You own architecture decisions across both SaaS and on-prem operating models, mentor engineers across teams, and drive multi-quarter ML infrastructure strategy. You are a force multiplier.What You'll Do:Define ML infrastructure architecture across the platform: model deployment strategy (vLLM, Triton, or TGI), pipeline engineering (MLflow or Kubeflow), and cloud-native infrastructure across major cloud platforms (AWS, Azure, or GCP)Set direction for ML system reliability: monitoring, latency / throughput / availability targets, and incident response across research and production environments.Mentor senior MLOps engineers; raise the operational bar across multiple teams.Drive cross-team initiatives that improve inference performance and cost-efficiency, including distributed training frameworks (DeepSpeed, FSDP, Accelerate).Partner with ML researchers, product, and engineering leadership on multi-quarter ML infrastructure strategy.Ensure ML infrastructure scales across managed SaaS and fully air-gapped on-prem deployments.What You'll Bring:10+ years of MLOps, ML infrastructure, or machine learning engineering with history of architectural ownership.Proven track record architecting large-scale model deployment (including LLMs) and ML infrastructure at scale.Deep cloud expertise across major cloud platforms (AWS, Azure, or GCP) and strong Python proficiencyMentorship record — engineers you have grown now operate independently at higher levels.Deep comfort architecting ML systems that run in both managed SaaS and on-premises / disconnected air-gapped environments.Kubernetes at architectural depth — GPU scheduling, multi-tenancy, operators, and the failure modes of distributed workloads on shared clusters.Strong communication, stakeholder management, and decision-making skills, with a passion for building diverse, inclusive engineering teams. Strong Preference:Ownership of production reliability at platform level: SLO definition, incident command, postmortem practice, and driving reliability improvements across teams rather than services.Architecture-level experience with distributed training and fine-tuning at scale (DeepSpeed, FSDP, Megatron-LM), including cluster design, checkpointing strategy, and failure recovery.Deep GPU systems knowledge: CUDA, NCCL, interconnect topology (NVLink, InfiniBand/RoCE), and diagnosing performance and communication problems across nodes.Model optimization strategy at portfolio level: quantization (FP8, AWQ, GPTQ), speculative decoding, with measurable cost or latency outcomes across multiple systems.Experience in regulated or security-constrained environments — compliance-driven architecture, model governance, lineage, audit, and secrets management.On-prem / air-gap ML delivery architecture experience at scale.Track record maturing MLOps practice in a growing organization: standards, platform abstractions, and paved paths that outlived your involvement.Bare-metal GPU cluster architecture, including scheduling (Slurm or Kubernetes) and hardware lifecycle in customer or owned data centers.Nice to Have:Conference speaking, technical writing, or industry thought leadership.Open-source contributions to inference, serving, or ML infrastructure projects, particularly maintainer-level involvement.C/C++ or CUDA kernel experience for performance-critical paths.Arabic language skills.Why AI71:Mission-Driven Work: Work on cutting-edge AI applications with a talented and passionate team, solving real-world challenges in critical sectors.Unparalleled Opportunity: This is a chance to innovate and solve real-world challenges using AI at a company with unique access to world-leading models and resources.Career Growth: We offer competitive compensation, benefits, and significant career growth opportunities as a foundational member of the team.World-Class Environment: Enjoy a flexible working environment and the latest tools & technologies needed to do your best work.
التعليقات
لا توجد تعليقات منشورة حتى الآن.
