Ambuj Logo

AMBUJ KUMAR TRIPATHI

AI Engineer
Loading Neural Weights...
Adaptive ReAct Portfolio | v2. Build 2026.07
Open to Roles

BuildingProduction-GradeAgentic AI Systems

I build production-grade GenAI systems that combine retrieval, reasoning, and real-world deployment across web and messaging platforms.

Explore Work
GeminiClaudeQwenDeepSeekLangGraphLangFuseRedisFastAPIDockerRenderMetaPineconeQdrantOpenRouterGoogle CloudIBM WatsonSupabaseReactNext.jsGeminiClaudeQwenDeepSeekLangGraphLangFuseRedisFastAPIDockerRenderMetaPineconeQdrantOpenRouterGoogle CloudIBM WatsonSupabaseReactNext.jsGeminiClaudeQwenDeepSeekLangGraphLangFuseRedisFastAPIDockerRenderMetaPineconeQdrantOpenRouterGoogle CloudIBM WatsonSupabaseReactNext.js
Live: 4 Services
99.790% Avg Uptime
296 ms Avg Latency
System Status ↗
Live AI Status
Adaptive ReAct 🟢 OnlineWeb Search ⚡ ReadyWhatsApp 📱 ConnectedOAuth 🔐 Authenticated

Full-Stack System Architecture

Full Stack System Architecture
LIVE LLM TELEMETRY

Production LLMOps & Observability

Production-grade observability powered by Langfuse with real-time tracing, latency analytics, cost optimization, and PII-safe telemetry from live AI deployments.

Core Performance Metrics

TTFT Metrics
Time to First Token (TTFT)
~1s latency ensuring fast start-up times for interactive streaming.
Max Latency
End-to-End Latency
Complex multi-step reasoning and retrieval bounded tightly within 5s.
Production Performance
Production Streaming Throughput
~250 tokens/sec sustained streaming during concurrent inference.
Langfuse Dashboard
Cost and Latency
P95 OPTIMIZATION

Unit Economics & Latency Optimization

  • Granular Token Tracking
  • Inference Cost Analytics
  • Edge Latency Insights
Langfuse Dashboard
PII Masking
PII SHIELD

Enterprise Data Privacy

  • Active PII Masking
  • GDPR Compliant
  • Zero Trust Pipeline
Langfuse Dashboard
Public Traces
LIVE TELEMETRY

Production Traces & Logging

  • Real-time Execution Tracking
  • Conversational Flows
  • Agentic Routing Logs
PRODUCTION TRACES
LATENCY ANALYTICS
PROMPT EVALUATION
Explore Production Architecture →

Featured Systems

Agentic RAG & MCP Orchestrator

Architected a LangGraph Router and FastMCP server. Autonomously fetches live RapidAPI (Stocks), GitHub analytics, and executes dynamic Gmail operations via a secure HITL SMTP bypass.

FastMCP ServerGmail HITLLive ToolingHigh Reliability
MCPLangGraphReactFastAPI
View Live System →

Adaptive ReAct Omnichannel RAG

A 10-Node LangGraph Agent featuring LLM Function Calling for live market data, Pinecone hybrid retrieval, WhatsApp Cloud integration, and an autonomous CRON-driven Headless Agent for daily AI newsletters.

10 NodesTool CallingCRON AgentHybrid Search99.790% Uptime---ms
LangGraphFastAPIPineconeQwen
View Live System →

Agentic Legal AI (Confidence-Gated HITL & Intent Routing)

Processed 31,500+ chunks across 20+ legal acts with Parent-Child Chunking. Secure OAuth 2.0 multi-tenant vector search on Qdrant.

31,500 Chunks3 Models5.6K DLs99.685% Uptime---ms
QdrantSupabaseReactOAuth
View Live System →

Citizen Safety AI Assistant

Deployed spaCy-driven trilingual routing with intelligent 112/100 emergency fallback signaling.

112 EmergencyspaCy3 Languages90.002% Uptime---ms
Llama 70BChromaDBspaCy
View Live System →

Professional Journey

Dec 2025 - Present

Independent AI Systems Builder

  • Architecting decoupled, production-grade agentic pipelines with FastAPI backends, enforcing strict PII masking and GDPR compliance for secure, enterprise-ready deployments.
  • Implementing advanced retrieval and custom document parsing strategies (Parent-Child Chunking, MRL) alongside Vision LLMs (VLMs) for multimodal data extraction.
  • Merging domain-specific fine-tuned models into complex RAG architectures to drastically improve contextual accuracy and response quality for specialized use cases.
  • Establishing rigorous modern observability frameworks to actively monitor and optimize p95/p99 latency metrics, ensuring optimal orchestration health across live systems.
Sep 2025 - Oct 2025

AI Prompt Engineer

Hogarth Worldwide (WPP Marketing) | Remote
  • Engineered strict constraint-based parameterization prompts (72-77 tokens) tailored for high-end commercial AI image generation systems (Flux, SDXL), ensuring absolute corporate brand compliance.
  • Developed the 'Smart AI Prompt Builder' pipeline tool utilizing advanced keyword classification mapping across 8 distinct business problem use-cases.
  • Executed adversarial red-teaming tests systematically to identify hidden hallucination edges and mitigate prompt instruction drift behaviors.
Jan 2022 - Aug 2024

Associate Engineer - Applied AI and Workflow Automation

British Telecom Global Services Pvt. Ltd. | Gurugram, India
  • Spearheaded Applied AI prototyping (starting 2023) by initially integrating OpenAI's APIs for natural language parsing, and progressively upgrading to the Gemini 1.5 Flash Vision API within Python automation pipelines. This automated the deep extraction of mission-critical parameters from dense, legacy network specifications.
  • Designed an internal natural language parsing mechanism coupled with an intuitive Gradio UI, allowing cross-functional operations teams to query real-time database deployments organically.
  • Drove critical FTTP infrastructure optimization by pioneering programmatic legacy conversion strategies on GIS frameworks that delivered a consistent 15% improvement across continuous network delivery life cycles.
  • Architected large-scale data manipulation pipeline modules using Python, standardizing ad-hoc data requests and reducing legacy manual efforts structurally by 60%.
Nov 2021 - Jan 2022

Operations and Maintenance Engineer

Tata Communications Transformation Services Ltd. | Lucknow, India
  • Optimized enterprise-scale backbone ISP infrastructure performance parameters utilizing comprehensive Orion network monitoring strategies to actively mitigate SLA breaches and trim total downtime metrics by 10%.
Dec 2017 - Jun 2018

Optical Fiber Execution Engineer

Annu Infra Construct India Pvt. Ltd. | Kolkata, India
  • Spearheaded classified subterranean optical routing deployments safely spanning 3 high-security state zones under the strict compliance mandates of the Ministry of Defense.
Oct 2013 - Nov 2014

System Administrator

Tata Consultancy Services (TCS) | Noida, India
  • Established comprehensive active monitoring heuristics on cloud production instances to deliver and maintain an impeccable 99% uptime availability benchmark.

Community Impact

🤗 Featured by Hugging Face👀 159K+ Views🌟 Open Source Contribution

Verified Industry Credentials

Click to View

Credential Portfolio
View Credentials →
+ 14 more verified certifications on detailed portfolio
0
Fine Tuned Models
0K+
Chunks
0K+
Downloads
0K+
Total Reddit Views
0
GitHub Stars
Core Architecture

Adaptive ReAct Systems

Designing multi-agent workflows with LangGraph that dynamically route tasks, search the web, and execute tools autonomously.

Real-time Web Search Fallback
Graph-based State Management
Automated Tool Selection
LLMOps

Production Ops

Metrics from live multi-agent deployment on Render & Vercel.

Uptime: 99.790%
Latency: 296ms
Incidents: 0 (30d)
Ecosystem

Engineering Stack

Modern tools for scalable AI.

LangGraph
FastAPI
Pinecone
Redis
Gemini
Qwen
React
Render
Open Source

Fine-Tuned Models

Custom quantized models hosted on HuggingFace for edge-deployment.

Format: GGUF (Q4_K_M)
Host: HuggingFace
Downloads: 5.6K+
Impact

Community Metrics

Driving open-source AI adoption in India.

Top 1% Poster on r/LangChainTop 1% Poster
(r/LangChain)
GitHub Stars: 65 ⭐
Forks: 22
Total Views (25 Posts): 230K+

Explore The Details

Engineering DocumentationSystem Architecture20+ ProjectsCase Studies
Explore Engineering DocsHow this works
20+ ProjectsArchitectureEngineering NotesDeployment Guide

Production Highlights

4 Public AI Systems
WhatsApp Integration
OAuth Authentication
Live Monitoring
Hybrid Retrieval
Real-time Web Search
Fine Tuned Models
Public APIs
Open Source Focus

Let's Build Something Great

Currently open for Agentic AI Architecture, ML Engineering, and specialized NLP roles. Drop a message or connect directly.

📍Gorakhpur, India
🌐Open to Remote / Hybrid
📧kumarambuj8@gmail.com
LinkedIn Profile
GitHub Repos
Usually responds within 24 hours
Preferences