YASH LOTHE
CARNEGIE MELLON UNIVERSITY
MS IN ARTIFICIAL INTELLIGENCE AND INNOVATION
OPEN TO MAY 2027 NEW GRAD AI/ML ROLESOPEN: MAY 2027

I build recommender and retrieval systems, the agents that use them, and the evaluations that catch them failing. Texas Instruments, Harvard, Ericsson, and a paper in PNAS Nexus.

I build recommender and retrieval systems, the agents that use them, and the evaluations that catch them failing.

02 / WORK

Selected work

Four systems, end to end: from the index, to the agents that use it, to the evals that keep them honest.

+18%RECALL@50

Two-Tower Recommender

End-to-end retrieval and ranking. A two-tower model learns dense user and item embeddings, FAISS handles candidate generation, and a PyTorch MLP reranks the shortlist. Beats matrix factorization and GBDT baselines on Recall@50, NDCG@10, and MRR.

PyTorchFAISSNumPyPandas

Logistic regression first, then GBDT. I only went two-tower once it was clear the baselines could not learn user and item representations jointly, which was the whole problem.

12/30OVER-REFUSALS SURFACED

Agentic LLM Red-Teaming

An eight-node LangGraph agent that takes a deployment spec, generates adversarial attacks, judges every response against a seven-verdict rubric, and clusters failures into families with mitigations. It treats over-refusal as a first-class failure, not just harmful compliance.

LangGraphLiteLLMPydanticStreamlit

At midpoint the planner was deterministic and the clusterer was a stub. The version that mattered came from making the picker LLM-guided, so it re-targets whichever attack surface just failed.

66.2%EM, HOP-ROUTED ENSEMBLE

SportsIQ Reasoning Agent

A ReAct agent over three domain-partitioned corpora (rules, tactics, game situations) with hybrid dense and sparse retrieval plus cross-encoder reranking, on the hardest tier of SportQA. Retrieval precision was 99.8%, and 91 of 96 wrong answers still had the right evidence retrieved, so the ceiling was reasoning calibration rather than knowledge.

Team project with Cady He and Siyani Vengatagiri.

FAISSBM25bge-reranker-v2-m3OpenAI Tools

SportRAG came first as a fixed decompose-then-retrieve pipeline. Handing tool choice to the model is what exposed the real pattern: every technique that helped multi-hop hurt single-hop.

100%MULTI-DIGIT ADDITION

LLaMA-Style Transformer From Scratch

Scaled dot-product and grouped-query attention, RoPE, SwiGLU, LayerNorm, top-p sampling, and AdamW, all implemented from scratch in PyTorch. Training and eval pipelines cover text generation, zero-shot sentiment on SST and CFIMDB, and arithmetic reasoning.

PyTorchRoPESwiGLUAdamW

Getting it working was the easy half. The ablations were the interesting part: finding the smallest architecture that still hits 100% on multi-digit addition tells you where model capacity actually goes.

03 / RESEARCH

Research

Peer-reviewed and preprint.

PEER-REVIEWEDIN PRESS

Like humans, language models demonstrate face-to-character biases

Steven A. Lehr, Yash Lothe, Mahzarin R. Banaji

PNAS Nexus, 2026

DOI 10.1093/pnasnexus/pgag247 · LIVE AUG 18
MANUSCRIPT · UNDERGRADUATE RESEARCH, BITS PILANI GOA
Deep Learning Approaches and Speech Pattern Analysis for Dementia Detection
PDF
04 / EXPERIENCE

Experience

Texas InstrumentsAI Solutions Engineer InternDallas TX · Jun 2026 to Aug 2026

Built per-account personalization for the production recommender behind TI.com using a Thompson sampling contextual bandit with three-tier fallback. Shipped a new email recommendation channel to production with a FastAPI serving endpoint, drift-gated testing, and MLflow deployment. Designed the RAG subsystem for an on-prem agent that auto-generates SPICE testbenches for analog circuits.

Harvard UniversityResearch Assistant, Banaji Implicit Social Cognition LabCambridge MA · Aug 2024 to Jul 2025

My undergraduate thesis, conducted at the lab for BITS credit. Led a responsible AI evaluation of large multimodal models, auditing face-to-character inference biases across four frontier models. Built the preprocessing, bias-quantification, and multi-model evaluation pipelines behind 8,160 trials.

EricssonAI/ML InternBangalore · Jun 2024 to Aug 2024

Designed synthetic data generation pipelines producing over 100,000 samples using GANs, VAEs, and kernel density estimation. Evaluated 15+ open-source LLMs on synthetic enterprise benchmarks with ROUGE, BERTScore, and COMET to guide internal model selection.

Goavega SoftwareSoftware Engineering InternBangalore · May 2023 to Jul 2023

Built and deployed a conversational AI agent with LangChain and the OpenAI API for a US client, enabling real-time auction data analysis. Delivered the full stack with Flask, PostgreSQL, and React, and presented the architecture and a live demo to the CEO, CTO, and COO.

Carnegie Mellon UniversityMS Artificial Intelligence and InnovationMay 2027
BITS Pilani GoaBE Computer ScienceMay 2025
05 / ABOUT

About

Yash Lothe, smiling, wearing a navy suit and light blue shirt, photographed outdoors in front of a hedge.

I've been captivated by ML and LLM agents since my junior year of undergrad and I have not really stopped thinking about it since. Most of what I build sits where retrieval meets evaluation: systems that find the right thing, and the harnesses that catch them when they don't. I care about that second half more than most people seem to. There are two sides to this technology, and the capability side already gets plenty of attention.

Outside of that: basketball, badminton, cricket, and an unreasonable number of movies. I like the ones that burn slow for an hour and then pull the rug out. I rapped in college. I do not anymore.

BASKETBALLBADMINTONCRICKETFILM
06 / CONTACT

Get in touch.

Open to new grad AI/ML engineering roles starting May 2027.