Anindya Mondal

I'm a final-year PhD researcher at the Surrey Institute for People-Centred AI (CVSSP), University of Surrey.

I am advised by Dr. Anjan Dutta, Dr. Xiatian Zhu, and Dr. Joaquin M. Prada, in collaboration with Dr. Sauradip Nag. I am also currently a Research Intern at Adobe.

I develop unified vision-language models that bridge visual understanding and controllable generation, spanning language-grounded perception, visual counting, and count-faithful synthesis.

Email  /  CV  /  Scholar  /  Github  /  LinkedIn

profile photo

News

  • Jul 2026: [New] Paper on unified understanding & generation accepted as a TOG Journal paper at SIGGRAPH ASIA 2026
  • Jun 2026: [New] Starting as a Research Intern at Adobe
  • Jan 2025: Paper on multi-label object counting accepted at AAAI 2025!
  • Dec 2024: Awarded AAAI 2025 Conference Travel Grant ($1,200)
  • Oct 2023: Presented actor-agnostic action recognition work at ICCV Workshop 2023 in Paris
  • Sep 2022: Started PhD at University of Surrey with full studentship funding

Work Experience

Adobe Research, Bengaluru, India
Position : Research Intern

Building rubric-aware VLM judges (reward models) to evaluate generated content.

Jun 2026 - Present
University of Surrey (CVSSP), Guildford, UK
Position : Doctoral Researcher

Research on unified vision-language models, generative AI, and RL post-training.

Oct 2022 - Present
Indian Institute of Science, Bengaluru, India
Position : Research Intern

Developed a source-free domain-adaptation framework for image classification.

May 2022 - Aug 2022
Jadavpur University, Kolkata, India
Position : Undergraduate Research Assistant

Graph learning and event-based vision for signal recovery and moving-object detection.

Oct 2020 - May 2022

Research

My research interests include computer vision, generative AI, and vision-language models. Some papers are highlighted.

ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation
Anindya Mondal, Sauradip Nag, Xiatian Zhu, Joaquin M. Prada, Anjan Dutta
SIGGRAPH ASIA (TOG), 2026
PDF / arXiv / code
TL;DR Single unified 3B-parameter VLM that jointly solves object counting, crowd counting, referring-expression counting, and count-faithful image generation via density-aware adaptive zooming, MLLM attention-derived objectness maps, and a cycle-consistent GRPO strategy with nested local, boundary, and global rewards, achieving state-of-the-art across seven benchmarks without any benchmark-specific training.
OmniCount: Multi-label Object Counting with Semantic-Geometric Priors
Anindya Mondal, Sauradip Nag, Xiatian Zhu, Joaquin M. Prada, Anjan Dutta
AAAI, 2025   (Oral Presentation)
DOI / arXiv / code
TL;DR Training-free multi-label counter coupling CLIP semantic embeddings for open-vocabulary category grounding with SAM point-prompt geometric priors for precise instance segmentation; introduces OmniCount-191, the first benchmark with point, bounding-box, and VQA multi-label count annotations.
CountLoop: Iterative Agent Guided High Instance Image Generation
Anindya Mondal, Ayan Banerjee, Sauradip Nag, Josep Llados, Xiatian Zhu, Anjan Dutta
Under Review
PDF / arXiv / Surrey
TL;DR Training-free diffusion pipeline with a VLM planner–critic loop: the planner generates structured instance layouts via chain-of-thought reasoning; the critic provides count and spatial feedback; instance-driven cross-attention masking with cumulative attention composition prevents semantic leakage across high-density scenes, reducing counting error by up to 57%.
Actor-agnostic Multi-label Action Recognition with Multi-modal Query
Anindya Mondal, Sauradip Nag, Joaquin M. Prada, Xiatian Zhu, Anjan Dutta
ICCVW, 2023
DOI / code
TL;DR Transformer decoder treating each action class as a multi-modal semantic query (CLIP visual + text embeddings), decoupling action classification from actor-specific pose topology for actor-agnostic multi-label recognition; state-of-the-art across five benchmarks spanning human and animal actions.
Time-varying Signals Recovery via Graph Neural Networks
John A. Castro-Correa, Jhony H. Giraldo, Anindya Mondal, Mohsen Badiey, Thierry Bouwmans, Fragkiskos D. Malliaros
ICASSP, 2023
PDF / DOI
TL;DR Encoder-decoder GNN (TimeGNN) trained with a composite loss of MSE and a Sobolev graph smoothness operator, exploiting spatio-temporal correlations across graph topology for robust missing-entry recovery on real sensor and traffic datasets.
Recovery of Missing Sensor Data by Reconstructing Time-varying Graph Signals
Anindya Mondal, Mayukhmali Das, Aditi Chatterjee, Palaniandavar Venkateswaran
EUSIPCO, 2022
PDF / DOI / code
TL;DR Formulates missing wireless sensor data recovery as time-varying graph signal reconstruction minimising a Sobolev norm combining graph Laplacian smoothness across spatial and temporal dimensions, surpassing state-of-the-art by up to 54% at high missing-data rates.
Moving Object Detection for Event-based Vision using Graph Spectral Clustering
Anindya Mondal, Shashant R, Jhony H. Giraldo, Thierry Bouwmans, Ananda S. Chowdhury
ICCVW, 2021
PDF / DOI / code
TL;DR Maps asynchronous event streams to a k-NN graph, applies graph Laplacian spectral decomposition for eigenvector-based clustering (GSCEventMOD), and determines the optimal cluster count via spectral gap analysis, enabling fully unsupervised moving object detection for event cameras.

Education

University of Surrey, United Kingdom
Position : Doctor of Philosophy (Ph.D.) in Artificial Intelligence

Thesis: Unified Vision-Language Models, Generative AI, RL Post-training and Object Counting

Oct 2022 - Late 2026 (Expected)
Jadavpur University, Kolkata, India
Position : Bachelor of Engineering (B.E. Hons.) in Electronics & Telecom

Thesis: Graph Signal Processing and Event-based Vision

Aug 2018 - May 2022

Miscellanea

Academic Service

  • Reviewer, CVPR
  • Reviewer, ICCV
  • Reviewer, ECCV

Teaching

  • Teaching Assistant, Computer Vision (University of Surrey)
  • Teaching Assistant, Machine Learning (University of Surrey)

Template from Jon Barron.