← WorkMundi · 1M+ jobs from around the world, liveSign inCreate free account

Senior AI/ML Solution Architect - Generative AI & Agentic Systems (Pune)

Flentas Technologies · Pune

📅 11/08/2026
🔔 Alert me about jobs like this
No password, no sign-up. Just the email — and you can leave the list anytime.
🔓 Apply — free →
Opens this job on WorkMundi. The account is free and takes under a minute.

See the other 113,540 jobs in India →

🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
About company: Flentas helps Startups, SMEs & Enterprises leverage the full potential of the cloud. We are AWS advanced consulting partners with certified cloud engineers to help you throughout every stage in your journey to Digital Transformation, helping you attain maximum ROI. Our services include cloud managed services, cloud-native application development, cloud migration, DevOps implementation, cloud governance automation, big data consulting, resource augmentation, IoT, and, Generative AI. Position Overview We are looking for a Senior AI/ML Solution Architect with deep expertise in Generative AI and agentic systems to lead the design and implementation of enterprise-scale AI solutions. This role requires a unique blend of hands-on technical expertise in both Large Language Models (LLMs) and Small Language Models (SLMs), combined with the architectural vision to deploy these solutions across diverse computing environments. The ideal candidate will architect scalable agentic solutions, implement advanced fine-tuning strategies, and design comprehensive integration systems that connect AI capabilities with enterprise applications. You will be at the forefront of our AI transformation initiatives, working with cutting-edge technologies while maintaining a practical approach to deployment and optimization. Experience Requirements - Overall Experience: 8+ years in technology and software development - Generative AI Experience: 2+ years of hands-on experience with LLMs and generative AI systems - Solution Architecture Experience: 4+ years architecting enterprise-scale solutions Key Responsibilities Architecture & Design - Design and architect scalable agentic solutions using advanced LLM capabilities - Implement Model Context Protocol (MCP) integrations to connect applications with diverse external services and APIs - Develop multi-agent orchestration systems for complex workflow automation - Design context and memory management systems for persistent agent interactions Technical Implementation - Build and optimize Retrieval-Augmented Generation (RAG) systems for efficient knowledge retrieval - Implement agent frameworks (LangChain, LangGraph, Semantic Kernel, Agno) for various deployment environments - Design and deploy model inference pipelines optimized for different computing environments (cloud, edge, on-premises) - Develop comprehensive fine-tuning strategies for both Large Language Models (LLMs) and Small Language Models (SLMs) - Architect SLM deployment strategies for resource-constrained environments - Implement model compression and quantization techniques for efficient inference Integration & Connectivity - Architect REST/gRPC/GraphQL APIs and SDK integrations for seamless service connectivity - Implement event-driven architectures using webhooks and message buses - Design secure authentication and authorization systems (SSO/OIDC) - Build connectors for popular platforms (Slack, Jira, Salesforce, CRM/ERP systems) Data & Model Management - Design comprehensive data preprocessing pipelines including cleaning, deduplication, and PII reduction - Implement embedding creation and re-embedding strategies for optimal retrieval - Develop chunking and windowing strategies for mobile-optimized content processing - Establish model selection criteria and evaluation frameworks Required Technical Skills Core AI/ML Expertise - Foundation Models: Deep experience with GPT-4, Claude, LLaMA, and other state-of-the-art LLMs - Small Language Models (SLMs): Expertise in deploying and optimizing SLMs (Phi-3, Gemma, TinyLlama) for mobile environments - Agent Frameworks: Proficiency in LangChain, LangGraph, Microsoft Semantic Kernel, Agno, and custom agent development - RAG Systems: Advanced knowledge of retrieval-augmented generation, vector databases, and semantic search Fine-tuning & Adaptation - Advanced fine-tuning techniques: LoRA/QLoRA, DoRA, AdaLoRA for parameter-efficient training - Model compression: Pruning, quantization (INT8/INT4), knowledge distillation - Prompt-tuning, adapters, prefix tuning, and P-tuning v2 methodologies - RLHF/RLAIF techniques for alignment and preference learning - Domain-specific fine-tuning for mobile use cases and vertical applications Deployment & Optimization - SLM Deployment: Expertise in deploying Small Language Models across various computing environments - Multi-Platform Optimization: Experience optimizing both LLMs and SLMs for cloud, edge, and on-premises deployment - Efficient Inference: Knowledge of quantization (GPTQ, AWQ, GGML), pruning, and distillation techniques - Model Compression: Advanced techniques for reducing model size while maintaining performance - Real-time Processing: Expertise in streaming inference and adaptive reasoning depth control - Performance Optimization: Proficiency in autoscaling, rate limiting, and resource management Adaptive Fine-tuning - Workplace-specific model adaptation and optimization - Federated learning .
Read the rest of the job →

Similar jobs

Job on WorkMundi — the world's largest job board. See more jobs from every continent, updated live.

📢
🎁

Before you apply, rehearse this interview.

Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card.

I want my training →