Research

My research focuses on understanding and improving how machine learning models organize complex information. My work spans interpretable vision-language learning, human-guided representation learning, and large-scale computational image analysis.

Vision-Language ModelsInteractive Machine LearningHuman-in-the-Loop ImagingComputational Imaging
Qualitative comparison of evidence patches for a woman riding a bike with a basket.
Qualitative result for a woman riding a bike with a basket: the visualization compares evidence patches selected by ItemizedCLIP and the prototype-based model for sentence- and object-level queries.
01Current Research

Interpretable Vision-Language Learning

Learning fine-grained and disentangled visual evidence for complex language queries.

Modern vision-language models can associate images with complex textual descriptions, but it is often unclear which visual evidence supports each individual semantic concept in the text. My current research investigates how fine-grained semantic structure can be explicitly introduced into vision-language representations to make their predictions more localized, disentangled, and interpretable.

In particular, I study prototype-mediated approaches that organize visual evidence through learnable semantic prototypes. Rather than relying solely on direct text-to-image attention, the model learns intermediate prototype structures that connect language concepts with relevant image regions without requiring dense segmentation annotations.

Research Questions

How can a model distinguish visual evidence for different concepts within the same description?

Can semantic structure emerge from weak supervision without pixel-level annotations?

Can prototype representations improve interpretability while preserving general vision-language capabilities?

Vision-Language ModelsPrototype LearningVisual GroundingWeak SupervisionInterpretable AI

Ongoing research · Manuscript in preparation

KeySI system overview connecting keyword curation, visualization, feedback translation, and model adaptation.
KeySI system overview: user-curated keywords and visual feedback are translated into document-level supervision for adapting text representations.
02Interactive Machine Learning

KeySI: Human-Guided Representation Learning

KeySI: An Interaction Framework for Tuning Text Embeddings Based on Human Feedback

Pretrained text embeddings capture general semantic relationships, but their representation of similarity may not align with the concepts that matter to a particular user. KeySI investigates how users can directly guide the organization of a learned embedding space through lightweight semantic feedback.

Users define and refine groups of meaningful keywords, which are translated into document-level supervision through retrieval and semantic filtering. The resulting feedback is then used to adapt the underlying representation model, allowing users to iteratively inspect and reshape the learned semantic space.

Key ideaTransform lightweight human semantic feedback into supervision for representation learning.
Interactive MLHuman-in-the-Loop AIRepresentation LearningVisual Analytics

IEEE Transactions on Visualization and Computer Graphics · IEEE VIS 2026

Coral CT imaging visualization for growth-band analysis.
CoralCT project image: CT and micro-CT imagery are used to support interactive analysis of coral skeletal growth bands.
03Ongoing Project

CoralCT: Human-in-the-Loop Analysis of Coral Growth from CT Imaging

Collaboration with Thomas M. DeCarlo · Tulane University

Coral skeletal growth bands provide valuable records of long-term coral growth and environmental change, but manually identifying and analyzing these structures across large CT datasets is time-consuming and requires domain expertise.

In collaboration with Thomas M. DeCarlo at Tulane University, I am working on machine-learning methods for analyzing coral skeletal CT and micro-CT imagery within the CoralCT platform. The project explores deep-learning approaches for detecting coral growth bands while incorporating expert feedback into the analysis process.

Human-in-the-loop workflowThe model proposes growth-band boundaries or candidate locations, experts review and correct the predictions, and those corrections can subsequently be used to improve the model.

An additional research question is how well these models generalize across coral species and imaging conditions, including the tradeoff between shared cross-species models and species-specific adaptation.

Coral CTMicro-CT ImagingGrowth Band DetectionExpert FeedbackModel Adaptation

Ongoing research

Developmental connectomics visualization of reconstructed cerebellar structures across mouse developmental stages.
Developmental connectomics visualizations from mouse cerebellum data, showing reconstructed neural structures and developmental-stage comparisons.
04Earlier Research

Computational Imaging for Developmental Connectomics

Harvard University · Lichtman Laboratory

At the Lichtman Laboratory at Harvard University, I worked on computational methods for analyzing large-scale electron microscopy data of the developing mouse cerebellum across multiple developmental stages.

My work focused on building and applying computational image-analysis methods for processing large microscopy datasets, including U-Net-based segmentation, feature-based image alignment, stitching-error detection, and reconstruction of biological structures across serial electron microscopy images.

Computational ImagingElectron MicroscopyU-Net3D ReconstructionConnectomics

Segmentation

Machine-learning-based segmentation of large-scale electron microscopy imagery.

Reconstruction

Image alignment, stitching, error detection, and serial-section reconstruction.

Analysis

Computational analysis of neuronal structures across developmental stages.