At ServiceNow's Now Assist Skill Kit team I work alongside PMs, ML engineers and applied research to design how humans curate, evaluate and trust agentic AI. Ground-truth anatomy, trajectory comparison, custom LLM metrics, LLM-as-Judge — the systems that decide what production AI is allowed to do.
Across eight years I've moved between government platforms, tax and healthcare, ERP and design systems — and the constant has been the same: heavy data, real constraints, and users who cannot afford a pretty guess.
Previously Intuit, Oracle, D. E. Shaw, Cerner and Philips. Master's from IDC, IIT Bombay, three published papers, and 8+ years across enterprise AI, SaaS, healthcare, ERP and government platforms.
Letting teams author their own LLM-judged metrics — defining criteria, scoring scales and rationale — so evaluation reflects the business, not just the benchmark.
Designing how teams author ground truth, compare agent trajectories against it, and read the result well enough to ship — or block — an AI agent.
A repeatable methodology for migrating sub-applications onto the Redwood design system — audit, IA, component mapping and rollout, deep-dived on one sub-application.
Construction of atomic lattices using marker-based AR, exploring modularity and inter-marker interaction. Presented at VRST '22, Tsukuba.
A tangible system that lets visually-impaired users perceive, create and modify line charts. India HCI 2020.
A tangible device that treats sound as an object you can sense — bound by physics, gravity and gesture. A study in playful, embodied UX.
Ground truth design, agentic trajectory comparison, LLM-as-Judge interfaces, evaluator workflows that production AI depends on.
Long-horizon thinking on permissions, data complexity, internationalization, and the workflows that run businesses — not just decorate them.
Building, governing and migrating component libraries — Oracle Redwood, and most recently the ServiceNow Horizon 2.0 system. Tokens, primitives, patterns.
Hardware-aware interaction design, AR/VR, and accessibility-first methods. The kind of craft that academic peer-review tends to reward.
Defining product vision with cross-functional leadership, leading 0→1 initiatives, and mentoring designers into stronger decision-makers.
Three published papers across VRST and India HCI. Comfortable in literature, lab studies, qualitative methods, and shipping the findings.
Evaluation harnesses, LLM-as-Judge prototypes and interface sketches for AI that has to be trusted — built outside the day job, where being wrong is cheap and fast.
Ask what's on the bench→What truly sets Anurag apart is his wisdom, his inquisitive nature, and his out-of-box mindset.
An excellent design leader, passionate about his craft and dedicated to bringing new ideas to the team.
A strategic mindset that lets him envision and execute long-term strategies that drive real outcomes.
He collaborates across cross-functional teams and adapts swiftly to evolving requirements.
Not just a talented designer — a great team player and a true tech enthusiast.
Consistently goes above and beyond to support his peers, demonstrating real empathy.
Career transitions across India and abroad. Mock interviews. Resume reviews. Long-term 1:1 mentorship. The conversations I wish I'd had at the start.
Book on Topmate →