# Siddharth Vohra > AI Research Engineer at AWS and M.S. Computer Vision candidate at Carnegie Mellon. Auditing where language and vision-language models fail when evidence is missing, misleading, or changed only in presentation. Also known as Sidd Vohra. Based in Pittsburgh, Pennsylvania. Site: https://siddvoh.com Hello, I'm Sidd! I'm an AI Research Engineer based in Pittsburgh, PA. I build AI-native products at Amazon Web Services and am pursuing a Master's in Computer Vision at Carnegie Mellon University's School of Computer Science (Robotics Institute). Before this, I built production ML systems for AWS Bedrock, Amazon Transcribe, and Amazon Translate, working across speech and language. I earned my Bachelor's in Mathematics-Computer Science from UC San Diego in 2022. ## Contact and profiles - Email: siddvoh@gmail.com - Google Scholar: https://scholar.google.com/citations?user=N9DDnEEAAAAJ - ORCID: https://orcid.org/0009-0002-6199-0485 - ResearchGate: https://www.researchgate.net/profile/Siddharth-Vohra-2 - GitHub: https://github.com/siddvoh - LinkedIn: https://www.linkedin.com/in/siddvoh ## Experience ### Amazon Web Services (Aug 2022 — Present) #### AI Research Engineer - Period: Sep 2025 — Present - Location: Pittsburgh, PA - Team: Generative AI Innovation Center - Note: Machine Learning Engineer until Aug 2026 Sole engineer on an AI-native engineering platform and multi-agent system that lets enterprise and public-sector teams create, deploy, govern and audit AI agents in one place. Designed a protocol from scratch that lets agents operate virtual machines inside AWS data centers, opening a secure and auditable path for agent deployment. Lead agentic AI projects for AWS partners and customers, running working sessions with C-suite executives at a European insurer and a US public-sector energy organization. #### Software Development Engineer II - Period: Apr 2024 — Sep 2025 - Location: Seattle, WA - Team: AWS AI · Bedrock Generative AI Services Feature lead across Amazon Transcribe, Bedrock Data Automation and Amazon Translate. Built speech-recognition models up to 30% faster and 2x more accurate on multilingual transcription, and served as technical lead for Amazon Translate across 30+ enterprise customers. #### Software Development Engineer - Period: Aug 2022 — Apr 2024 - Location: Seattle, WA - Team: Lambda · App Runner · Elastic Beanstalk - Related: Using WAF with App Runner in Copilot, AWS Copilot CLI Blog — https://aws.github.io/copilot-cli/blogs/apprunner-waf/ Led the AWS WAF integration into the Copilot CLI so App Runner applications ship secure by default, built an automated testing suite for the App Runner console, and helped launch App Runner in three new regions. ### Teradata (Jul 2021 — Sep 2021) #### Software Engineer Intern - Period: Jul 2021 — Sep 2021 - Location: San Diego, CA Restructured large-object storage in the TeraCloud architecture, improving computation efficiency by roughly 50%. ## Publications Peer-reviewed and accepted work only. Nothing under review is listed. ### Does the Selected Object Reach the Reader? Auditing Identity Handoffs in Grounded Language-Model Pipelines - Venue: EMNLP 2026, Grounding Language Models Workshop - Date: 29 October 2026, Budapest, Hungary - Authors: Siddharth Vohra, Runmin Jiang, Xiaomo Li, Min Xu - Link: https://arxiv.org/abs/2609.04579 Traces whether an item selected early in a grounded QA pipeline survives retrieval and reaches the answer model. On 1,463 aligned HybridQA records, body-only BM25 dropped it out of the top five 26.6% of the time; hybrid reranking cut that to 1.0%. ### Research Agents Feed Doubts, Not Beliefs: Auditing Belief Framing in Live-Web Search - Venue: EMNLP 2026, Research on Agent Language Models Workshop - Date: 29 October 2026, Budapest, Hungary - Authors: Siddharth Vohra, Min Xu Four web-search agents tested on 48 expert-checked health claims across 1,874 runs. Stating doubt shifted every system toward rejection and more than doubled searches seeking disconfirming evidence. Replay tests traced roughly three quarters of the shift to how agents read the evidence they retrieved. ### The Audit Decides the Verdict: Instrument Effects Rival Demographic Bias in LLM Decision Audits - Venue: EMNLP 2026, Research on Agent Language Models Workshop - Date: 29 October 2026, Budapest, Hungary - Authors: Siddharth Vohra, Manikandan Ravikiran - Link: https://arxiv.org/abs/2609.09048 Across 34,070 decisions from five models in hiring, lending and medical triage, none of 36 planned demographic comparisons survived multiple-comparison correction, while listing order and a single instruction rewrite moved outcomes more than any measured demographic effect. ### When Models Defer to Wrong Answers: A Robustness Audit of Source-Attributed Cues in Multiple-Choice QA - Venue: EMNLP 2026, Grounding Language Models Workshop - Date: 29 October 2026, Budapest, Hungary - Authors: Manikandan Ravikiran, Siddharth Vohra - Link: https://arxiv.org/abs/2609.08934 Across 220,000 responses from four models in English, Hindi, Bengali, Tamil and Telugu, an unverified expert cue pulled models to a named wrong option in 41.1% of trials they had first answered correctly. A majority cue did so in 12.5%. ### Absent-Byte Diagnoses: Auditing Structured Medical VLM Interfaces - Venue: MICCAI 2026, Agentic AI for Medicine Workshop - Date: 1 October 2026, Strasbourg, France - Authors: Siddharth Vohra, Manikandan Ravikiran - Link: https://siddvoh.com/hearsay/ Three of five medical vision-language models returned a structured diagnosis in 617 of 3,000 calls where the prompt claimed an image was attached but the request carried no image bytes. A client-side verifier blocked all 5,047 tested evidence-binding violations. ### Hearsay: Vision-Language Medical Diagnoses Without an Image - Venue: ICMR 2026, Toward Trustworthy Vision-Language Models in the Wild Workshop - Date: 2026, Amsterdam, The Netherlands - Authors: Siddharth Vohra - Link: https://siddvoh.com/hearsay/ Audits frontier vision-language models under missing-image medical prompts and identifies structured diagnostic confabulation, demographic sensitivity, and structured-output failure modes. ### TEEMIL: Towards Educational MCQ Difficulty Estimation in Indic Languages - Venue: COLING 2025, Main conference - Date: January 2025, Abu Dhabi, United Arab Emirates - Authors: Manikandan Ravikiran, Siddharth Vohra, Rahul Verma, Rohit Saluja, Arnav Bhavsar - Link: https://aclanthology.org/2025.coling-main.142/ Introduces the TEEMIL-H and TEEMIL-K datasets for Hindi and Kannada MCQ difficulty estimation, with multilingual model baselines and ablations over context, answer options and none-of-the-above options. ### You Reap What You Sow — Revisiting Intra-class Variations and Seed Selection in Temporal Ensembling for Image Classification - Venue: COMSYS 2023, Springer - Date: 2023, Singapore - Authors: Manikandan Ravikiran, Siddharth Vohra, Yuichi Nonaka, Sharath Kumar, Snehanshu Sen, Nidhi Mariyasagayam, Krishnan Banerjee - Link: https://doi.org/10.1007/978-981-19-0105-8_8 Studies how intra-class variability, seed size and seed selection affect semi-supervised Temporal Ensembling performance across image-classification datasets. Ravikiran and Vohra contributed equally; names are ordered alphabetically. ### Investigating the Effect of Intraclass Variability in Temporal Ensembling - Venue: arXiv 2020, Preprint - Date: 21 August 2020 - Authors: Siddharth Vohra, Manikandan Ravikiran - Link: https://arxiv.org/abs/2008.08956 Early study of how within-class variation and the number and choice of labelled seeds affect Temporal Ensembling. Led to the 2023 Springer paper above. ## Press ### Carnegie Mellon University, 20 July 2026 "Healthcare Blind Spots: AI models prone to fabricating diagnoses" - Link: https://www.ri.cmu.edu/healthcare-blind-spots-ai-models-prone-to-fabricating-diagnoses/ Featured audit of frontier medical vision-language models uncovering severe diagnostic confabulation and evidence-binding failures when images are detached or missing. ### The National, 27 July 2026 "Patients warned off using AI chatbots for self-diagnosis as flaws revealed" - Link: https://www.thenationalnews.com/news/uae/2026/07/27/patients-warned-off-using-ai-chatbots-for-self-diagnosis-as-flaws-revealed/ Broad international reporting examining clinical safety risks and the necessity of client-side verification in commercial healthcare agent deployments. ## Academic service ### Conference peer review Reviewer and programme committee member at NeurIPS, ICML, EMNLP, MICCAI and ACM KDD, covering mechanistic interpretability, biomedical retrieval, and the robustness and safety of agentic AI systems. - NeurIPS 2026: Ethics Review, main conference; Symmetry and Geometry in Neural Representations; AI for the Global South; Trustworthy AI for Good; AI for Meta-Science: Scaling and Organizing Science in the Age of AI Scientists; AI for Drug Discovery: Bridging the Translation Gap; Interpretability for Discovery: Understanding and Discovering Novel Knowledge in AI Models; Grounded and Faithful Vision-Language Models for Real-World Deployment; Continual Learning for Enterprise AI Agents; Dynamics at the Frontiers of Learning, Sampling and Games; AI and the Self: Human Identity, Authenticity and Agency in the Age of AI - EMNLP 2026: Grounding Language Models: Learning Faithfully and Efficiently; Workshop for Research on Agent Language Models - COLM 2026: Context Beyond the Window: Persistent Knowledge in Language Models; Social Simulation with LLMs: Fidelity in Applications; Workshop on Efficient Reasoning - ECCV 2026: Women in Computer Vision; Computer Vision for Ecology - ICML 2026: Compositional Learning: Safety, Interpretability and Agents - MICCAI 2026: Workshop on AI for Safe Surgery - ACM KDD 2026: Programme Committee; Agentic Software Engineering (SE 3.0): The Rise of AI Teammates - IJCAI-ECAI 2026: Workshop on AI for the Global South - ACM SIGCITE 2026: Main conference ### Hackathon judging Invited to the judging panels of flagship university hackathons at Carnegie Mellon, UC San Diego and Cincinnati, deciding awards across competitive fields of several hundred teams. - HackCMU 2025, ACM@CMU, Carnegie Mellon University, September 2025. 600+ participants, 100+ submissions. - DiamondHacks 2025, ACM at UC San Diego, April 2025. 650+ registrations, 95 projects. - RevolutionUC 2025, ACM@UC, University of Cincinnati, March 2025. 873 applications, 300 selected. ## Recognition and memberships ### Member, Pittsburgh Section, IEEE (Member #96729838) Member in good standing of the Institute of Electrical and Electronics Engineers, valid through December 2026. ### Professional Member, ACM (Member since May 2026) Admitted to the Association for Computing Machinery having fulfilled the requirements for Professional Membership. ### Fellow, Guild of Expert Engineers, Hackathon Raptors Association (August 2026) Elected by peer review. Election requires at least four of five votes from randomly selected sitting Fellows against eight published criteria. - Link: https://www.raptors.dev/our-members ### Gemini Academic Program Award, Google (2026) $20,000 in Google Cloud credits for Gemini foundation-model research and multimodal evaluation audits. ## Education ### M.S. Computer Vision, Carnegie Mellon University (Expected Dec 2026) School of Computer Science, Robotics Institute. ### B.S. Mathematics-Computer Science, University of California, San Diego (2019 — 2022) Cum laude, Provost Honors all quarters. Founding and principal member of the Machine Learning Club. ## Interactive research pages - Hearsay: vision-language medical diagnoses without an image: https://siddvoh.com/hearsay/ - Identity handoffs in grounded language-model pipelines: https://siddvoh.com/identity-handoffs/