Research

Computational imaging and inverse problems; diffusion and generative priors for reconstruction.

Computational Imaging Inverse Problems Generative Priors Computational Photography

I'm interested in sensor–algorithm co-design: how to shape the measurement (exposure, sampling, readout) so that a generative prior has an easy job at reconstruction time. Recent work uses diffusion priors to recover full-resolution images and HDR radiance from a single, heavily constrained capture.

Peiran Li, Fei Chen, Xixin Wu
Interspeech 2025
(a) A rolling-shutter camera exposes rows with a coded long/medium/short exposure-time map, producing one RAW coded measurement. (b) A hash-grid MLP predicts HDR radiance, which is pushed through the same forward model and fit to the measurement. (c) Compared with the best single exposure, both the +3 EV shadows and the −6 EV highlights are recovered.
Single-Shot HDR Imaging with Exposure Control
Ongoing
Dartmouth College · Prof. Adithya Pediredla
May 2026 – Present · Research Assistant

Instead of bracketing three exposures, expose different rows of a single frame for different durations, then recover the full HDR radiance from that one coded measurement.

  • Formulated the sampling operator as row-wise long/medium/short exposures over a differentiable RAW sensor model, and captured a coded HDR dataset on an ArduCam module.
  • Recovered the HDR radiance from a single coded measurement with a pretrained generative prior, optimized self-supervised under an SNR-weighted loss.
  • Swapped in posterior-sampling (DPS, DAPS) and MAP/variational (RED-diff, plug-and-play) priors under the same forward model to compare them.
(a) Stratified random sampling directly on the sensor lattice cuts readout bandwidth. (b) A fine-tuned SD3 prior reconstructs the full image from a 6.25% pixel readout. (c) Example results against ground truth.
Towards Bandwidth-Limited Imaging with Generative Priors
Under review
Dartmouth College · Prof. Adithya Pediredla & Prof. Sotiris Nousias
Aug 2025 – Present · Research Assistant

Sensor readout, not pixel count, is the bottleneck in modern imaging. Rather than binning or cropping, read out a random 6.25% of pixels at native resolution so aliasing becomes incoherent, then let a diffusion prior fill in the rest.

  • Sampled 6.25% of pixels with a stratified random mask at native sensor resolution, turning the coherent aliasing of uniform subsampling into noise-like artifacts a prior can remove.
  • Completed the full image with an SD3 diffusion backbone, conditioning a trained ControlNet on the mask and a nearest-neighbor interpolation of the samples.
  • Outperformed downsample-then-super-resolve pipelines (Swin2SR, HAT, StableSR, Stable Diffusion ×4 Upscaler) at matched budget: LPIPS 0.192 vs 0.294, FID 10.99 vs 13.83, and every metric on Urban100.
  • Built a galvo-mirror / APD single-pixel scanner to emulate sparse readout optically along Lissajous trajectories, and validated the pipeline on self-captured RAW data.
EEG-based Speech Decoding with Multi-mode Joint Modeling
Interspeech 2025
Centre for Perceptual and Interactive Intelligence, Hong Kong · Prof. Xixin Wu
Dec 2024 – Aug 2025 · Research Intern

Imagined speech is the hardest EEG mode to decode. Training jointly on imagined, intended, and spoken speech gives the model a much stronger signal than imagined speech alone.

  • Collected EEG from 11 participants across three speech modes: imagined, intended, and spoken.
  • Jointly modelled the three modes, augmenting EEGNet with dynamic masking over temporal-spatial features to reach 34.95% four-way vowel accuracy on imagined speech.
Gaze-Based Weak Supervision for Medical Image Segmentation
Medical imaging
The Chinese University of Hong Kong · Prof. Qi Dou
Sep 2024 – Nov 2024 · Research Intern

Where a clinician looks is a cheap, dense annotation. Eye-tracking traces can replace most of the pixel-level labeling for segmentation.

  • Developed a weak-supervision scheme combining gaze annotation with multi-level learning from discriminative attention; operated eye trackers and preprocessed gaze with CRFs.
  • Cut annotation time from 12 hours to 2.2 hours with no loss in segmentation accuracy.

Most of my projects start at the optical table. I like building the capture rig myself so the forward model in the paper matches the one on the bench.

Single-pixel scanning setup on an optical table: an avalanche photodiode (APD), a lens, a Thorlabs galvo mirror, bright lights, and a checkerboard scene.
The single-pixel scanner used for the bandwidth-limited imaging project. A galvo mirror steers the line of sight through a lens onto a single APD, so the sensor reads exactly the pixels the Lissajous trajectory visits on the checkerboard scene.
Galvo / APD scanning
Built and aligned a galvanometer-mirror scanner with a single-pixel avalanche photodiode, driving the mirrors along resonant Lissajous trajectories to realize sparse sampling optically.
SPAD laser ranging
Time-of-flight depth acquisition with single-photon avalanche diodes, including laser–detector synchronization and photon-timing histograms.
EEG & eye tracking
Operated EEG recording systems and eye trackers, and ran the full data-collection protocol with human participants.
Slide
Open PDF ← → to navigate · Esc to close