Talari Pradeep

AI Infrastructure Engineer |

I architect scalable AI infrastructure, platform MLOps/LLMOps pipelines, and multi-agent orchestration frameworks — from model design to production rollout. Passionate about self-healing clusters, speed, and high reliability.

0 Years Exp.
0 Projects
0 Certs
0 % Uptime
sre-shell@tp-core:~
LIVE
visitor@tp-shell:~$

The Engineer Behind the Stack

Talari Pradeep
Bangalore, India
Open to Work

I'm an AI Infrastructure & Platform Engineer with a passion for building robust, scalable platforms that power intelligent applications.

With 6+ years of IT operations experience and expertise across Azure, AWS, Kubernetes, and Terraform, I specialize in designing LLMOps pipelines, multi-agent orchestration frameworks (LangGraph), and semantic document search engines (RAG) that run at scale.

Working as an Infrastructure & AI Engineer at Accenture in Bangalore, I focus deeply on platform reliability, self-healing architectures, and automated MLOps pipelines — enabling teams to transition models to production seamlessly.

Multi-Agent Orchestration & LangGraph OS
Cloud-Native MLOps/LLMOps on AKS & EKS
Enterprise RAG Platforms & Vector DBs
Full-Stack Observability & OpenTelemetry

AI Infrastructure & Platform Sandbox

Experiment with live AI infrastructure parameters, trace network flows, scale model worker nodes, and audit platform resources.

⏱️

SLA & Error Budget

Configure a Service Level Agreement (SLA) to calculate allowed downtime budgets, and test your reaction speed when chaos hits.

Target SLA: 99.9%
99.0% 99.9% 99.99% 99.999%
Weekly Budget 1.68 hrs
Monthly Budget 7.31 hrs
Yearly Budget 3.65 days
System: OPERATIONAL
Error Budget Remaining: 100%

ArgoCD GitOps Rollout

Simulate a GitOps continuous deployment pipeline. Push a commit and watch ArgoCD synchronize container pods in real time.

💻
Git Repo
v1.2.0
ArgoCD
Synced
Ingress Route: us-east-1 (Primary)

us-east-1 (Primary) Online

pod-0
pod-1
pod-2
pod-3

eu-central-1 (DR) Standby

pod-dr-0
pod-dr-1
pod-dr-2
pod-dr-3

Chaos & Auto-Healing

Inject live infrastructure faults into the metrics dashboard and observe the autonomous SRE control loop restore operations.

CPU Usage
28%
RAM Usage
45%
API Latency
120ms
Error Rate
0.0%
SRE Automation Log (Live Stream)
[HEALER] Health check active - all probes OK.
📖

SRE Runbook Simulator

Execute interactive checklists to triage outages, resolve state drifts, and roll back canary versions.

Select a runbook scenario to start the SRE checklist.

Let's Connect

Open to full-time roles, freelance contracts, and interesting collaborations. Drop me a message!