
About this role
The candidate will be responsible for building and evaluating GenAI services, including RAG and agentic reasoning systems. They will work closely with Frontend and Backend teams to integrate AI agents into real products with a focus on reliability and safety. The technical environment involves working with frameworks like LangChain, PyTorch, and Hugging Face, and deploying LLM API integrations and self-hosted models like Llama 3. The role involves managing the AI model lifecycle, monitoring, and implementing guardrails to ensure safety and and quality. The engineer will work on projects involving vector databases, semantic search, and cloud platforms like AWS, GCP, or Azure.
Skills & technologies
Must have
Nice to have
- Automated testing for hallucination rates
- Human-in-the-loop mechanisms
- Human-in-the-loop
- Memory systems design
- Prompt Engineering
- Context Engineering
Read full description
Hard Skills:
- Practical experience developing and evaluating GenAI services, including RAG systems, understanding of the "ReAct" (Reasoning + Acting) paradigm and Agentic RAG, LLM API integration, and prompt/context engineering, as well as training, fine-tuning, and deploying ML models in production environments.
- Knowledge of ML/GenAI frameworks: LangGraph or LangChain, PyTorch / TensorFlow, Hugging Face, OpenAI/Anthropic SDK.
- Practical commercial experience working with AI models via API (Gemini, Anthropic) and self-hosted models (Llama 3, Mistral, Mixtral), with an understanding of model differences based on functional/non-functional requirements (FR/NFR).
- Deep understanding of how LLMs interact with external APIs via Function Calling.
- Experience with vector databases, semantic search methods, and principles of database structuring and cleaning.
- Practical experience with at least one cloud platform (AWS, GCP, or Azure).
- Proficiency in Python and understanding of asynchronous programming.
- Understanding of the AI model lifecycle: monitoring, versioning, and quality evaluation (RAGAS, DeepEval), with hands-on experience using these tools.
- Experience with Guardrails: setting hard constraints on conversation topics and agent actions, filters that automatically strip personal data before sending requests to external LLMs, and the ability to build output filters that fact-check generated responses before they're displayed.
- Deterministic Logic Integration — running AI agents on strict schemas to prevent the model from "making things up."
- A plus: knowledge of automated testing approaches for evaluating responses across large datasets to measure hallucination rates before MVP launch.
- Knowledge of Human-in-the-loop mechanisms, ensuring agents cannot execute actions without final user verification.
- Ability to design memory systems that store context from a client's previous conversations and operations for personalization (Long-term Memory & User Context).
- Ability to clearly communicate complex technical concepts and mentor team members.
- Ability to quickly test and evaluate new libraries and approaches.
- Analytical problem solving - debugging complex "black boxes" and understanding why an agent behaves unpredictably.
- Focus on delivering a working prototype rather than a perfect research paper.
- Close collaboration with Frontend and Backend developers to seamlessly integrate AI agents into the required environment.
- Flexible working format - remote, office-based or flexible
- A competitive salary and good compensation package
- Personalized career growth
- Professional development tools (mentorship program, tech talks and trainings, centers of excellence, and more)
- Active tech communities with regular knowledge sharing
- Education reimbursement
- Memorable anniversary presents
- Corporate events and team buildings
- Other location-specific benefits
- not applicable for freelancers