I work at the intersection of product judgment and AI systems: ambiguous problem spaces, human-AI interfaces, and decisions that need to be trustworthy, not just fast.
An AI-assisted workspace that helps finance teams investigate and resolve the small percentage of transactions that automated reconciliation cannot match — combining AI-generated explanations, evidence, and human review to reduce month-end closing time and improve accuracy.
An AI assistant that helps product managers generate hypotheses, prioritize experiments, define success metrics, and analyze results — making experimentation faster and more structured without requiring deep analytics expertise.
An AI-powered payroll operations platform that automates document processing, payroll validation, anomaly detection, and employee issue resolution — reducing manual effort while improving payroll accuracy and compliance.
An AI-powered vendor document review and verification workspace that helps operations teams classify submitted documents, extract key information, catch missing or inconsistent details, and resolve exceptions before activating a vendor.
A mobile-first citizen grievance platform for Andhra Pradesh, built for citizens the official system leaves out — no Aadhaar login required, Telugu-first, file a complaint from any phone in under three minutes.
How should a PM decide whether an AI feature is reliable enough to ship?
Compares three extraction configurations — zero-shot, few-shot, and structured schema with a self-check pass — on accuracy, hallucination rate, cost per run, latency, and how much still needs a human at an adjustable review threshold.
Finding — the mid-cost config matched the most expensive one's accuracy at under half the cost and latency. The pricier option didn't automatically win.
At what AI confidence level should a decision auto-approve versus require human review?
120 AI-screened expense decisions, split across auto-approve / analyst review / manual investigation by two adjustable thresholds — with automation rate, missed issues, financial exposure, and false positives recalculating live wherever you set them.
Finding — a $2,140 fraudulent claim scored 94% confidence by the model. High confidence and being right aren't the same thing — the default threshold lets it through with zero human review.
Rajesh works across AI product strategy, systems design, and execution — mostly in the space where a model's output has to become someone's actual decision. His strongest work happens before the roadmap hardens: clarifying what the system should be confident about, identifying the few decisions that matter, and shaping the interface so people trust it enough to act.
How confidence, evidence, approval states, and reversibility turn model outputs into decisions people can safely act on.
Takeaway — People do not trust AI because it sounds intelligent. They trust it when they can understand its recommendation, challenge it, and recover when it is wrong.
Why adding an approval button is not enough—and how AI products should define confidence thresholds, ownership, escalation, and accountability.
Takeaway — Human oversight works only when the system makes it clear what requires review, when intervention is necessary, and who owns the final decision.
How to communicate commitments, experiments, and unresolved assumptions without pretending the future has already been decided.
Takeaway — A useful roadmap shows what is committed, what is being tested, and what still needs evidence.