التخصص
تكنولوجيا المعلومات
نوع الدوام
دوام كامل
الخبرة
أي مستوى
الموقع
أبوظبي
وصف الوظيفة
Associate MLOps Engineer On Site | Abu Dhabi AppliedAI is a pioneering AI technology company headquartered in Abu Dhabi, UAE. We are committed to innovation and excellence in artificial intelligence solutions across regulated industries such as healthcare, insurance, government, and financial services. AppliedAI is at the forefront of redefining the future of work through cutting-edge AI solutions. We empower organizations by automating complex, document-heavy processes, ensuring unparalleled efficiency and accuracy. Our commitment to synergizing human intelligence with artificial intelligence helps our clients excel in highly regulated industries, including healthcare, finance, and insurance. We're seeking an Associate ML Ops Engineer to join our growing team. In this role, you'll work at the intersection of operations and development, supporting our platform's reliability, performance, and security while helping our development teams build and maintain scalable solutions. This is a strong opportunity for someone early in their SRE/MLOps career to build technical depth alongside a senior team. Key ResponsibilitiesHelp monitor and maintain production, development and staging environments, supporting high availability and performance of our architectureCollaborate with DevOps, MLOps and Development teams to help troubleshoot and resolve issuesSupport our observability stack, with guidance from senior engineersHelp maintain and improve our incident response processesAssist with capacity planning and performance optimization for systems handling up to 10,000 requests per minuteSupport compliance with security standards and regulatory requirements across our infrastructureParticipate in on-call rotation during regular business hours, with senior engineer backup for escalationsHelp implement and maintain SLOs, SLIs, and SLAsSupport cloud cost optimization efforts alongside system performance and reliabilitySupport our CI/CD pipelines and deployment processesRequired Skills & Experience Infrastructure Best PracticesFamiliarity with infrastructure best practices such as the AWS and Azure Well-Architected FrameworkBasic understanding of infrastructure patterns for high availability, fault tolerance, and disaster recoveryUnderstanding of infrastructure security fundamentals (principle of least privilege, network segmentation, encryption at rest/in transit)Awareness of infrastructure compliance and governance frameworksInterest in cost optimization strategies and FinOps practicesExposure to Infrastructure as Code concepts (modularity, reusability, versioning)Understanding of observability patterns (logging, metrics, tracing)Technical Experience1-2 years of experience in SRE, DevOps, or a similar role (internships and hands-on project experience will be considered)Some experience with AWS services, including:- Compute: Lambda, ECS, Fargate - Networking: ALB, ELB, API Gateway, Route53, CloudFront, AppSync - Databases: DynamoDB, RDS (PostgreSQL), Aurora - Messaging: EventBridge, SNS, SQS - Security: Security Groups, Secrets Manager (SM), Systems Manager (SSM), IAM - Developer Tools: ECR, CodeBuild, CodeDeployExposure to monitoring and observability toolsBasic knowledge of infrastructure as code (CDK and/or Terraform)Understanding of event-driven architecture conceptsSome experience with containerization and microservicesBasic scripting and automation skillsGood problem-solving abilities and a systematic approach to debuggingExperience working in Agile environments is a plusPreferred Qualifications:AWS / Azure / GCP certificationsExperience with Next.js, Node.js and PythonFamiliarity with authentication systems (Auth0, SSO)Awareness of regulatory compliance requirements (SOC 2, HIPAA, GDPR, PCI DSS)Interest in ML/LLM operationsExposure to multi-region AWS deploymentsInterest in working with high-traffic systemsBasic experience with database management and optimizationAwareness of caching strategies and CDN implementationsUnderstanding of data lifecycle management and ETL processesExposure to vector and graph databases is a plusWhat We Offer:Opportunity to work with cutting-edge technologiesMentorship and guidance from senior SRE, Architecture, DevOps and MLOps engineers as you build your skillsCollaborative environment with dedicated Architecture, Development, DevOps & MLOps teamsGrowth potential in a rapidly scaling startupWork with a globally distributed teamRegular working hours with flexibility for occasional emergency supportA clear path to take on more ownership of SRE practices as you grow in the roleRequired Qualities:Strong communication skillsProblem-solving mindsetTeam player attitudeSelf-motivated and proactive approachAbility to work in a fast-paced startup environmentStrong interest in continuous learning and skill developmentThe ideal candidate is early in their reliability engineering career, eager to build technical depth in a collaborative environment, and comfortable working in a dynamic startup setting while developing toward higher standards of system reliability and performance over time. Benefits:Opportunity to work with a leading AI technology company.Collaborative and innovative work environment.Growing, entrepreneurial and forward-thinking culture. Career growth and professional development opportunities.Exposure to a thriving ecosystem working from our Abu Dhabi HQ.21 days of paid annual leave.Company provided health insurance.Visa sponsorship for international candidates.
التعليقات
لا توجد تعليقات منشورة حتى الآن.
