Research

Trustworthy AI for human work.

I study how AI-assisted systems can be evaluated for reliability, especially when speech, generated documentation, and human decision-making meet.

StatusOngoing
RecognitionUROP funded
Funding periodsSpring + Summer 2026
InstitutionBoston University
01

Research statement

Reliability is part of the interface.

AI systems used for documentation should do more than generate fluent text. They should help people recognize uncertainty, potential errors, and unsupported content without adding unnecessary friction to the work.

My interests span trustworthy AI, medical documentation, speech and language processing, DSP, and human-centered evaluation. The common question is simple: how do we know a system is helping, and how do we make its failure modes visible?

02

Current project

AI-assisted clinical documentation

This ongoing project investigates an end-to-end prototype that combines automatic speech recognition, transcript processing, medical-document generation, and validation mechanisms intended to flag potential errors or unsupported outputs.

01 Clinical conversation Input context
02 Automatic speech recognition Audio → transcript
03 Transcript processing Structure + context
04 Document generation Draft documentation
05 Validation mechanisms Flag possible issues
06 User study + evaluation Human interaction
Abstract research architecture. No private data or confidential implementation details are shown.
03

My role

Prototype, validation, and study design.

  • Developing components of the end-to-end prototype and validation layer.
  • Investigating ways to surface potential errors and unsupported outputs.
  • Designing a user-study and data-collection workflow around human interaction and documentation impact.
  • Preparing ongoing work for a future research manuscript without representing it as published or accepted.
04

Research questions

What needs to be measured?

  • Which generated statements are unsupported by the available source context?
  • How should a validation layer communicate risk without overwhelming the user?
  • How does validation affect review behavior, documentation quality, and time?
  • Where do automated metrics stop being useful, and where is human evaluation essential?
05

Evaluation approach

Separate system quality from human impact.

The evaluation plan distinguishes component-level behavior from end-to-end use. Prototype testing examines transcription, generation, and validation behavior; the planned user workflow evaluates how people interact with the system and how it affects documentation work.

01Component behavior
02Unsupported-output review
03Human-centered outcomes
06

Boundaries

What is intentionally not shown.

This public case study contains no protected health information, private datasets, confidential implementation details, or clinical performance claims. Adviser, laboratory, poster, report, and manuscript links will be added only after they are confirmed and public.