AI Engineer
Explicitly requires vibe coding skills; mentions Cursor, GitHub Copilot, and Claude Code to accelerate engineering.
About the Role
As an AI Engineer at VenPep Solutions, you will design, build, and deploy domain-specific MCP servers, end-to-end RAG pipelines, and scalable vector-search infrastructure to enable production-grade agentic AI systems. You will integrate and manage LLMs, vector databases, cloud AI platforms, and mentor junior engineers while promoting AI-assisted development workflows.
Job Description
Role
The AI Engineer will design, build, and operate production-ready AI systems focused on Model Context Protocol (MCP) servers, retrieval-augmented generation (RAG) pipelines, and agentic AI workflows. The role involves integrating LLMs, vector databases, cloud AI platforms, and AI-assisted development tools while mentoring junior engineers and improving system accuracy.
Key Responsibilities
- Design, build, and deploy context-specific MCP (Model Context Protocol) servers tailored to business domains, including defining tool schemas, resource endpoints, and prompt templates
- Architect and manage end-to-end RAG pipelines: document ingestion, chunking strategies, embedding generation, vector store management, and retrieval optimization
- Select, configure, and manage vector databases for semantic search and knowledge retrieval at scale
- Design agentic AI workflows using frameworks such as LangChain, LlamaIndex, AutoGen, or CrewAI, orchestrating multi-step reasoning and tool usage
- Implement and maintain LLM integrations (OpenAI, Anthropic Claude, LLaMA, Mistral), including prompt engineering, context window management, and fine-tuning pipelines
- Champion vibe coding practices and leverage AI-assisted development tools (Cursor, GitHub Copilot, Claude Code) to increase engineering velocity
- Build scalable AI infrastructure on cloud platforms (AWS Bedrock, Azure OpenAI Service, GCP Vertex AI)
- Conduct model evaluation, retrieval quality benchmarking, and continuous improvement of AI system accuracy
- Mentor junior AI engineers on MCP architecture, RAG best practices, and responsible AI development
- Stay current with AI research and translate emerging techniques into production-ready solutions
Requirements
- 5 to 7 years of experience in AI/ML engineering with a focus on applied LLM and generative AI
- Hands-on experience designing and deploying MCP servers and defining tool schemas and prompt templates
- Deep expertise in RAG system design: chunking strategies, embedding models (OpenAI, Cohere, sentence-transformers), retrieval pipelines, re-ranking, and hybrid search
- Strong proficiency in Python and LLM orchestration frameworks (LangChain, LlamaIndex, or equivalent)
- Experience with vector databases: Pinecone, Weaviate, ChromaDB, Qdrant, or pgvector
- Proficiency with OpenAI API, Anthropic API, Azure OpenAI Service, or AWS Bedrock
- Familiarity with vibe coding workflows and AI-assisted development tools (Cursor, GitHub Copilot, Claude Code)
- Experience deploying AI services to production using Docker, Kubernetes, or serverless technologies (AWS Lambda, Azure Functions)
- Strong understanding of prompt engineering, system prompt design, and context management techniques
- Excellent research, problem-solving, and technical communication skills
Nice to have
- Contributions to open-source MCP server implementations or AI agent frameworks
- Experience with fine-tuning LLMs: LoRA, QLoRA, or full fine-tuning on domain-specific datasets
- Knowledge of real-time inference optimization: model quantization, ONNX, vLLM, or TGI
- Familiarity with evaluation frameworks for RAG quality: RAGAS, TruLens, or DeepEval
- Experience with agentic AI platforms such as AutoGen, CrewAI, or custom multi-agent orchestration
Location
In person