Shreya Mendi

MEng AI Student @ Duke | ML Engineer | AI Product Builder

|

About

About Shreya

Here is a little background

Focused on building intelligent systems that create real-world impact. I'm an MEng AI student at Duke University with experience in ML engineering, DevOps, and AI product development. I enjoy turning research ideas into practical solutions β€” whether that's a safety-critical CV system for search and rescue, a bias audit for transportation AI, or a multimodal fashion discovery engine. I bring a strong problem-solving mindset and love working at the intersection of AI and social good.

β˜• Fueled by chai tea, not coffee
πŸ”΅ Duke Blue through and through
πŸ€– Believes AI should be explainable
🌿 Building tech for social good
πŸŽ“ Duke University β€” MEng in Artificial Intelligence

Experience

Duke Trust Lab

Research Assistant

Duke Trust Lab

PythonPyTorch

Jan 2026 β€” Present

  • Co-authoring a NeurIPS submission on safe agentic AI behavior and intervention policy under faculty supervision; leading experimental design, evaluation framework construction, and manuscript preparation.
  • Developed benchmarking pipelines across 3+ research projects measuring model calibration, AI transparency, and decision reliability in high-stakes deployment settings.
BMW Group

Student Consultant (AI Capstone)

BMW Group

PythonFastAPI

Jan 2026 β€” Present

  • Built and evaluated multiple ML regression models (LightGBM, Ridge, MLP, Optuna-tuned TabularMLP) on real BMW dealer sales data to predict inventory turnover speed; ranked product configurations by predicted sell performance across the full vehicle lineup.
  • Delivered explainable AI-driven inventory recommendations through a live REST API and interactive dealer dashboard, enabling business teams to make data-informed spec decisions with transparent model confidence.
Duke University

Teaching Assistant β€” Managing AI in Business

Duke University

Jan 2026 β€” Present

  • Facilitate discussion sections and office hours for a graduate-level course covering AI strategy, LLM deployment, and responsible AI adoption; supporting 40+ MEng students on project design and technical implementation.
  • Evaluate student AI system proposals and provide structured feedback on model selection, deployment tradeoffs, and business impact framing across industry-sponsored capstone projects.
Assetmantle

DevOps Engineer

Assetmantle

AWSKubernetesDocker

Sep 2023 β€” May 2025

  • Reduced AWS/Hetzner infrastructure costs by 38% through architecture optimization and automated CI/CD (Docker, Kubernetes) with rollout/rollback policies, maintaining 99%+ uptime across distributed nodes.
  • Streamlined CI/CD pipelines with Kubernetes which increased deployment frequency and reduced time-to-market across the engineering org.
Hewlett Packard Enterprise

Software Development Intern

Hewlett Packard Enterprise

DockerJenkinsPythonLinux

Jan 2023 β€” Jul 2023

  • Built Dockerized services on Linux, automating CI/CD workflows with Jenkins, shell scripting, and cloud computing platforms.
  • Integrated REST APIs in Python for monitoring using Grafana & Prometheus improving the overall stability of the product.

Skills

Hover over a skill for current proficiency

Python

95%

Python

PyTorch

85%

PyTorch

TensorFlow

82%

TensorFlow

FastAPI

88%

FastAPI

Docker

85%

Docker

Kubernetes

80%

Kubernetes

AWS

90%

AWS

Google Cloud

75%

Google Cloud

SQL

85%

SQL

Git

90%

Git

Jenkins

75%

Jenkins

Linux

85%

Linux

Pandas

90%

Pandas

NumPy

90%

NumPy

JavaScript

78%

JavaScript

TypeScript

72%

TypeScript

React

70%

React

Next.js

68%

Next.js

Jupyter

92%

Jupyter

MLflow

80%

MLflow

Projects

35 projects Β· hover to explore

When2Speak β€” LLM Intervention Policy Agent
AI ResearchNLP

When2Speak β€” LLM Intervention Policy Agent

Trained a lightweight RL policy network for multi-agent dialogue intervention using PyTorch and NLP. Reduced unnecessary interventions by 25% while maintaining task success rate, validated on a 10,000-dialogue simulation suite via A/B testing. Co-authoring a NeurIPS submission on safe agentic AI behavior under faculty supervision at Duke Trust Lab.

PythonPyTorchReinforcement Learning
UAV-SAR β€” Aerial Human Detection for Search & Rescue
Computer VisionDeep Learning

UAV-SAR β€” Aerial Human Detection for Search & Rescue

Fine-tuned Faster R-CNN on thermal SAR imagery with domain-specific augmentations (snow, smoke, sensor noise). Achieved 20% recall improvement under adverse conditions with <5% clean-data accuracy loss. Built end-to-end CV pipeline in PyTorch with custom data loaders and metrics optimized for safety-critical deployment.

PythonPyTorchTensorFlow
BMW Capstone β€” Industrial AI Inventory Decision System
MLIndustry

BMW Capstone β€” Industrial AI Inventory Decision System

Built and evaluated ML regression models (LightGBM, Ridge, MLP, Optuna-tuned TabularMLP) on real BMW dealer sales data to predict inventory turnover speed. Delivered explainable AI-driven inventory recommendations through a live REST API and interactive dealer dashboard, enabling data-informed spec decisions without requiring ML expertise.

PythonFastAPIDocker
Safe-T β€” AI Equity Audit for Transportation Safety
AI EthicsData Analysis

Safe-T β€” AI Equity Audit for Transportation Safety

Audited transportation AI allocation algorithms for racial and economic bias across Durham, NC β€” documenting that Black residents (32% of population) account for 47% of pedestrian/cyclist crash victims, driven by systematic undercounting of demand in low-income and high-minority census tracts. Built an interactive Leaflet mapping platform comparing AI-predicted vs. need-based infrastructure allocation using Census, NCDOT crash records, and OpenStreetMap data.

PythonJupyterJavaScript
Mirror β€” Reflective AI Mental Wellness Companion
RAGMental Health

Mirror β€” Reflective AI Mental Wellness Companion

Built a reflective AI journaling companion using RAG over 8 psychological frameworks (CBT, IFS, NVC, Attachment Theory) to pattern-match emotional entries. Implemented cognitive distortion profiling across 4 distortion types and weekly self-awareness reports. Deployed full-stack Python/FastAPI on Railway with session persistence, mood trend tracking (1–10 daily score), and a trigger map logging recurring emotional patterns.

PythonFastAPIDocker
CineStyle β€” Multimodal Fashion Discovery from Film & TV
Multimodal AIRecommendation Systems

CineStyle β€” Multimodal Fashion Discovery from Film & TV

Built a film-to-fashion identification platform using FashionCLIP (512-dim embeddings) and FAISS GPU vector search over 20,000 Fashionpedia garment crops across 46 categories. Four-stage recommendation pipeline (FAISS β†’ NeuMF β†’ SASRec β†’ diversity filter) with measurable NDCG@10 gains over a popularity baseline. Deployed FastAPI on Railway, Next.js on Vercel; evaluated on Precision@K, Recall@K, and MAP@K with 500 synthetic users Γ— 30 interactions.

PythonNext.jsFastAPIDocker
Inflationship β€” Macroeconomic Forecasting Pipeline
Time SeriesForecasting

Inflationship β€” Macroeconomic Forecasting Pipeline

Built a SARIMAX inflation forecasting pipeline combining port traffic indicators with CPI data, reducing forecast error to 0.67–1.69% MAPE across major CPI categories. Validated model stability via 5-fold rolling cross-validation, achieving reliable predictive lift over CPI-only baselines.

PythonPandasNumPy
Alba β€” AI Carbon Footprint Tracker
SustainabilityChrome Extension

Alba β€” AI Carbon Footprint Tracker

Built a real-time, privacy-first Chrome extension estimating energy, carbon, and water footprint for AI prompts using model metadata and GitHub Models API. Implemented live footprint labels, heuristic + AI prompt optimization, and a client-side dashboard with daily impact summaries.

JavaScriptChrome ExtensionGitHub API
AI Audit β€” EU AI Act Compliance System
AI ComplianceMLOps

AI Audit β€” EU AI Act Compliance System

Built an EU AI Act compliance assessment system using TF-IDF + Logistic Regression, rule-based article evaluation (Articles 5, 6, 9, 10, 14), and automated remediation planning. Deployed full-stack ML workflow with MLflow, FastAPI, Docker, and Streamlit on Google Cloud Run, enabling explainable risk scoring and documentation audits.

PythonDockerFastAPIGoogle Cloud
Semantic Jury β€” Legal Semantic Search Engine
NLPVector Search

Semantic Jury β€” Legal Semantic Search Engine

Developed a semantic search engine for legal documents using embedding-based retrieval and citation-link modeling to surface semantically similar case law and statutes. Implemented vector search pipelines with passage ranking to support explainable legal research and faster knowledge discovery.

PythonTensorFlowJupyter
Edie Cursor
HTML

Edie Cursor

Edie is an AI cursor that fills a clinical form with you while you talk to your patient. Page and demo video: https://shreya-mendi.github.io/edie-cursor/ - Edie's cursor (light blue) follows the clinician-patient conversation. She moves to the question being discussed and leaves a suggested answer on it, with the exact words she heard.

PrecedentVector
HTML

PrecedentVector

How did approved gene therapies prove it, and what will FDA ask of you? 28 licensed gene therapies Β· 235 FDA documents Β· 5,766 pages Β· every answer cites a page number You are designing the safety and efficacy package for a gene therapy. You want to know what the approved products in your class actually did, and what FDA made them do

Website Portfolio
CSS

Website Portfolio

πŸ”— Live site: shreya-mendi.github.io/Website-Portfolio Personal portfolio site showcasing my projects and experience as an AI engineer. Built with plain HTML, CSS, and JavaScript, hosted on GitHub Pages.

Youandme
TypeScript

Youandme

πŸ”— Live app: youandme-xi.vercel.app Meals, workouts & dates β€” together. u-n-me is a private, two-person shared life app for a couple: a weekly meal plan, a collaborative grocery note, a workout planner, a date-night planner, a lightweight nudge feed, a recipe editor, and real recommenders that learn what the household likes. It has a tiny serverless backend (Vercel

TypeScript
Unmasking Reflections
HTML

Unmasking Reflections

πŸ”— Live site: shreya-mendi.github.io/UnmaskingReflections/ My reflections on the Unmasking AI book

Mindguard Benchmark
Python

Mindguard Benchmark

πŸ”— Live site: shreya-mendi.github.io/mindguard-benchmark/ > Content Warning: This repository contains synthetic prompts simulating mental health crises at various severity levels. The content is designed for AI safety evaluation research only. If you or someone you know is in crisis, please contact the 988 Suicide & Crisis Lifeline (call or text 988) or the Crisis Text Line (text HOME to 741741).

Python
Epstein Paper Trail
Python

Epstein Paper Trail

πŸ”— Live demo: shreya-mendi.github.io/epstein-paper-trail/ > "The Documents Don't Lie" Paper Trail is an NLP system that classifies legal consequences for individuals named in the Epstein case using entirely public documents. It features a RAG-based chatbot, a named entity recognition (NER) pipeline, consequence classification across four tiers, and a dark interactive timeline UI.

Python
Tradecraft
Python

Tradecraft

πŸ”— Live demo: shreya-mendi.github.io/Tradecraft/ > Five specialized AI agents that research, signal, risk-check, execute, and audit trades β€” with a polished dark-luxury dashboard. β”œβ”€β”€ server.py # FastAPI backend (Phase 2) β”œβ”€β”€ orchestrator.py # CLI pipeline runner (Phase 1) β”œβ”€β”€ requirements.txt

Python
Catanist
Python

Catanist

πŸ”— Live demo: shreya-mendi.github.io/Catanist/ An instrumented arena where LLM players compete at Settlers of Catan β€” the sister project to the Mafia Arena. Different models (via GitHub Models) sit around one board, each primed with a controlled intent (diplomatic, cutthroat, deceptive, greedy, …), and you spectate the match in a cute illustrated replay

Python
Mafia Arena
Python

Mafia Arena

πŸ”— Live demo: shreya-mendi.github.io/Mafia-arena/ An instrumented arena where LLM players play social-deduction Mafia. Unlike a win-rate leaderboard, this treats persona, intent (cruel, peace-mongering, deceptive, truth-seeking, …), group size, and composition as controlled knobs β€” then records everything and renders a watchable replay.

Python
Why2speak Labelling
Python

Why2speak Labelling

A friendlier version of the CSV: each labeller reads a model's reasoning trace and picks which kind of intervention the reasoning claims to make (6 labels). The rubric is always on-screen, progress autosaves, and it's blind to the GPT-5 judge. pip install -r requirements.txt streamlit run app.py Opens at http://localhost:8501. Each person enters a name; their labels save to

Python
PAEwall
Python

PAEwall

Multimodal, faithfulness-grounded patent infringement discovery. Given a U.S. patent, PAEwall ranks the companies most likely infringing it, generates a per-limitation claim chart scored against real product documentation, and co-generates red-team non-infringement and invalidity arguments β€” giving patent holders and defendants the same honest view of enforcement viability.

Python
Eco City
Python

Eco City

This project builds a 3D interactive eco-city simulation where a Reinforcement Learning (RL) agent learns to design and manage a sustainable city over time. Unlike static simulations, the agent: - makes sequential urban planning decisions - balances competing objectives (growth vs sustainability) - adapts to dynamic environmental feedback

Python
PoolCue Assist
Python

PoolCue Assist

AIPI 590 | Midterm Project Platform: Raspberry Pi 4 | Language: Python 3 Pool Cue Assist teaches correct billiards stroke mechanics by detecting and classifying cue stroke quality in real time. The goal is stroke straightness and smoothness β€” the most common cause of missed shots is cue twist and lateral wobble on the follow-through, not aim. The device uses an IMU mounted on the cue to measure ho

Python
Wordle XAI
Python

Wordle XAI

> MEng AI Final Project β€” Explainable Artificial Intelligence > A multimodal agentic AI that plays Wordle by fusing a vision model and an > information-theoretic NLP solver, then surfaces real-time explanations of > why it made each decision β€” bridging the gap between raw model predictions > and human-understandable reasoning.

Python
Style2Fit

Style2Fit

A project by Shreya Mendi.

Sharky
Python

Sharky

A Python project by Shreya Mendi.

Python
BMW Optimizer
Jupyter Notebook

BMW Optimizer

A Jupyter Notebook project by Shreya Mendi.

Jupyter
ReinforcementLearning
Jupyter Notebook

ReinforcementLearning

A project by Shreya Mendi.

Jupyter
QuietSky

QuietSky

QuietSky is a calming, encouragement-focused speech adventure game designed to help players practice fluent, confident speech through fun gameplayβ€”not through clinical or corrective cues The game contains four distinct modes, each modeled after genres the player enjoys (Flappy Bird, light exploration, light combat, and cockpit/FPS HUDs). Each mode supports a different aspect of fluent speech behav

XAI
Jupyter Notebook

XAI

A project by Shreya Mendi.

Jupyter
DISNEYYYY
Jupyter Notebook

DISNEYYYY

We used ChatGPT 5 on 11/11 at 12:45pm to generate some code and content for this project. This repo contains a compact machine-learning pipeline that predicts whether a Disney hero headlines a blockbuster (above-median inflation-adjusted gross) by blending: - Character descriptors (hero, villain, signature song) from data/disney-characters.csv.

Jupyter
ATM Modelling
Jupyter Notebook

ATM Modelling

Modelling the ATM data for alternate data class For the data.ipynb script to work.Select all files in the atm folder and Download the data from duke.box.com . It will automatically be a .zip file. The python notebook is optimized to use the .zip folder and extract all the datasets into dataframes.

Jupyter
Weather Modelling
Jupyter Notebook

Weather Modelling

Temperature prediction using RDU data for modelling

Jupyter
Sourcing Happiness
Jupyter Notebook

Sourcing Happiness

This project analyzes the World Happiness Report data to explore how happiness scores evolve across countries and regions over time (2019–2024). We investigate the role of key explanatory factors such as GDP, social support, healthy life expectancy, freedom of choice, generosity, and perceptions of corruption.

Jupyter

Contact

I've got just what you need. Let's talk.