Charlotte Vision Lab

The Charlotte Vision Lab is a computer vision research group at UNC Charlotte focused on building systems that understand and act in the visual world. Our work spans video understanding, multimodal learning, robotic perception, generative modeling, and trustworthy machine vision. Our goal is to build trustworthy systems that can perceive, reason, and assist in complex real-world environments.

Our Research Directions

Video Understanding for Activities of Daily Living

Temporal action detection, dense activity recognition, and long-video reasoning over unedited, real-world sequences.

Vision-Language and Vision-Language-Action Models

Visual question answering, domain adaptation, interpretable decision-making, and embodied reasoning.

Reliable 3D and Generative Vision

3D scene understanding, controllable image generation, and uncertainty estimation in open-world settings.

Highlights

ECCV 2026

VisCoP

Visual probing for efficient domain adaptation of vision-language models with minimal changes to pretrained parameters.

View paper

ECCV 2026

Ego2Exo VLM

Learning egocentric cues from exocentric video to improve vision-language understanding of daily living activities.

View paper

ECCV 2026

3D FaceShell

A privacy-preserving framework that controls how vision-language models interpret 3D face avatars while preserving identity.

View paper

CVPR 2026

MS-Temba

A Multi-scale Mamba-based Temporal Modeling Framework for efficient temporal action detection in long untrimmed videos.

View paper