Applied AI Engineer|

Rajan N

I build production LLM systems — RAG pipelines, multi-cloud model orchestration, and agentic architectures — and I break them on purpose too, through independent AI-safety and prompt-injection research.

01 — About

Grounded in shipping, curious about breaking things.

I'm an AI engineer who owns systems end-to-end — from architecture and model selection through deployment, monitoring, and the on-call reality of keeping an LLM system correct under real constraints: latency, cost, and trust. My work spans retrieval-augmented generation, LLM inference optimization, and multi-cloud model orchestration across AWS Bedrock, GCP Vertex AI, and Azure OpenAI, shipped and operated in production, not left in a notebook.

Outside of building, I run independent security research on LLM applications — probing production chat systems for prompt-injection and policy-bypass vulnerabilities, and reporting findings through coordinated disclosure. I think the two disciplines sharpen each other: you design better guardrails, and take deployment more seriously, once you've tried hard to break what you shipped.

0 Year building production LLM systems
0 Cloud LLM platforms orchestrated
0 Independent research projects shipped

What I'd bring to your team

  • I ship LLM features end-to-end — architecture, model choice, evaluation, and deployment — so a team spends less time handing work across roles and more time shipping.
  • I design for correctness under real constraints, not demo constraints: citation grounding, temporal validity, and hallucination rate as release-blocking metrics, not afterthoughts — see the numbers in the NBFC Compliance Intelligence case study.
  • I bring an adversarial mindset to what I build, from independently finding and responsibly disclosing an LLM safety vulnerability — the same instinct that catches failure modes before a customer does.
  • I default to measuring, not asserting: every claim on this site links to the evidence behind it.

02 — Skills

Toolbox

Retrieval-Augmented Generation (RAG)✓ Retrieval Engineering✓ LLM Orchestration & Model Routing LoRA / QLoRA Fine-Tuning✓ Prompt & Context Engineering LLM Application Security✓ LLM Benchmarking✓ Agentic Workflow Design Guardrails & Safety Evaluation✓ Vector Search & Hybrid Retrieval Data Engineering Pipelines Containerized Deployment & Orchestration Classical ML & Deep Learning Natural Language Processing Embeddings & Chunking Strategy✓ Retrieval Metrics (recall@k, MRR)✓ LLM Observability & Tracing✓ Synthetic Data Generation for Fine-Tuning✓ Inference Optimization & Cost Control Vector Databases & Caching LLM Response & Prompt Caching Transformer Architectures Tokenization & Text Preprocessing Cloud Infrastructure Design (AWS EC2, RDS, S3) GPU Infrastructure for Model Fine-Tuning Cloud IAM & Access Control Cloud Cost Management (FinOps) Production API Design (FastAPI) SSO & Identity Federation Cryptographic Hashing & Data Integrity (SHA-256) Asynchronous Messaging & Queueing (SNS, SQS) Secrets Management (AWS Secrets Manager)

03 — Featured Work

Selected Projects

Full-Stack · AI for Regulated Finance Live Alpha

NBFC Compliance Intelligence

Retrospective RBI-compliance auditor for Indian NBFC lenders, deployed and actively iterating

An AI-assisted compliance engine that ingests a lender's own loan documents and collections-call transcripts, extracts structured facts, resolves which RBI regulations were legally in force on the date of the event, and issues a cited, auditable verdict — compliant, violation, ambiguous, or no-clause-found — with every citation validated against a versioned regulatory corpus before it's ever shown to a user.

Four independently-testable stages (extract → retrieve → decide → validate), zero raw-document persistence by architectural constraint, row-level tenant isolation, and a production incident I diagnosed and fixed myself: sequential per-fact LLM calls were blowing past reasonable timeouts, fixed by fanning them out concurrently for a ~4x wall-clock improvement, with a regression test so it can't silently regress.

FastAPI + Celery/Redis Postgres 16 + pgvector LiteLLM Fly.io + Neon

Status: live alpha — pipeline deployed end-to-end; trust properties met (0 hallucinated citations, 100% deterministic-rule accuracy); judgment-based accuracy still weak and actively being improved. Numbers on the case study are live and honest, not polished.

Independent Security Research

AI Safety Vulnerability Disclosure

Responsible disclosure to a major LLM chat platform

Identified and responsibly disclosed a prompt-injection vulnerability affecting safety-policy enforcement in a major LLM chat product. Submitted a coordinated disclosure report to the vendor's security team, including reproducible test cases, supporting evidence, a technical writeup, severity assessment, and mitigation recommendations.

LLM application security Prompt-injection testing AI safety evaluation Coordinated disclosure

Status: submitted, currently in the vendor's disclosure review process

Read the full case study →

04 — Contact

Let's talk.

Open to Applied AI Engineer / LLM systems roles. Reach out directly — I reply fast.