Anindya Mondal

I am a final-year Ph.D. researcher at the Centre for Vision, Speech and Signal Processing (CVSSP) and the Surrey Institute for People-Centred AI, University of Surrey, advised by Dr. Anjan Dutta, Dr. Xiatian Zhu, and Dr. Joaquin M. Prada, while working closely with Dr. Sauradip Nag. Concurrently, I am a Research Scientist Intern at Adobe Research.

My research focuses on establishing verifiable spatial reasoning, compositional generation, and self-correcting capabilities in multimodal foundation models. I have authored 7+ peer-reviewed papers in premier venues (including ACM TOG/SIGGRAPH Asia, AAAI, and ICCV). My agenda centers on three core directions:

  • RL Alignment & Post-Training: Developing rubric-aware Process Reward Models (PRMs) and cycle-consistent GRPO objectives to align multimodal policies and suppress generative hallucinations.
  • Inference-Time Agentic Loops: Designing training-free MLLM planner–critic architectures that scale test-time compute to verify and steer diffusion dynamics toward compositional correctness.
  • Fine-Grained Spatial Grounding: Bridging discrete structural reasoning and continuous visual synthesis through dense semantic-geometric priors (e.g., CLIP + SAM).

Beyond research, I am committed to academic mentorship, having supervised 10+ MSc dissertations and served as a Teaching Assistant for Master’s-level Computer Vision and Machine Learning courses.

Email  /  CV  /  Scholar  /  Github  /  LinkedIn

profile photo
🚀 Looking for Opportunities: I am actively looking for Postdoctoral Associate or Research Scientist positions in generative modeling, agentic systems, and multimodal understanding areas. Please feel free to reach out here.

News

  • Aug 2026: [New] Paper on agent guided image generation accepted at TMLR
  • Jul 2026: [New] Paper on unified understanding & generation accepted as a TOG Journal paper at SIGGRAPH ASIA 2026
  • Jun 2026: [New] Starting as a Research Intern at Adobe
  • Jan 2025: Paper on multi-label object counting accepted at AAAI 2025!
  • Dec 2024: Awarded AAAI 2025 Conference Travel Grant ($1,200)
  • Oct 2023: Presented actor-agnostic action recognition work at ICCV Workshop 2023 in Paris
  • Sep 2022: Started PhD at University of Surrey with full studentship funding

Work Experience

Adobe Research, Bengaluru, India
Position : Research Intern

Building rubric-aware VLM judges (reward models) to evaluate generated content.

Jun 2026 - Present
University of Surrey (CVSSP), Guildford, UK
Position : Doctoral Researcher

Research on unified vision-language models, generative AI, and RL post-training.

Oct 2022 - Present
Indian Institute of Science, Bengaluru, India
Position : Research Intern

Developed a source-free domain-adaptation framework for image classification.

May 2022 - Aug 2022
Jadavpur University, Kolkata, India
Position : Undergraduate Research Assistant

Graph learning and event-based vision for signal recovery and moving-object detection.

Oct 2020 - May 2022

Research

My research interests include computer vision, generative AI, and vision-language models. Some papers are highlighted.

ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation
Anindya Mondal, Sauradip Nag, Anjan Dutta
ACM Transactions on Graphics/ SIGGRAPH ASIA, 2026
PDF / arXiv / web
TL;DR Single unified 3B-parameter VLM that jointly solves object counting, crowd counting, referring-expression counting, and count-faithful image generation via density-aware adaptive zooming, MLLM attention-derived objectness maps, and a cycle-consistent GRPO strategy with nested local, boundary, and global rewards, achieving state-of-the-art across seven benchmarks without any benchmark-specific training.
CountLoop: Iterative Agent Guided High Instance Image Generation
Anindya Mondal, Ayan Banerjee, Sauradip Nag, Josep Llados, Xiatian Zhu, Anjan Dutta
TMLR, 2026  
PDF / arxiv / Surrey
TL;DR Training-free diffusion pipeline with a VLM planner–critic loop: the planner generates structured instance layouts via chain-of-thought reasoning; the critic provides count and spatial feedback; instance-driven cross-attention masking with cumulative attention composition prevents semantic leakage across high-density scenes, reducing counting error by up to 57%.
OmniCount: Multi-label Object Counting with Semantic-Geometric Priors
Anindya Mondal, Sauradip Nag, Xiatian Zhu, Anjan Dutta
AAAI, 2025  
DOI / arXiv / web
TL;DR Training-free multi-label counter coupling CLIP semantic embeddings for open-vocabulary category grounding with SAM point-prompt geometric priors for precise instance segmentation; introduces OmniCount-191, the first benchmark with point, bounding-box, and VQA multi-label count annotations.
Actor-agnostic Multi-label Action Recognition with Multi-modal Query
Anindya Mondal, Sauradip Nag, Joaquin M. Prada, Xiatian Zhu, Anjan Dutta
ICCVW, 2023
DOI / code
TL;DR Transformer decoder treating each action class as a multi-modal semantic query (CLIP visual + text embeddings), decoupling action classification from actor-specific pose topology for actor-agnostic multi-label recognition; state-of-the-art across five benchmarks spanning human and animal actions.
Time-varying Signals Recovery via Graph Neural Networks
John A. Castro-Correa, Jhony H. Giraldo, Anindya Mondal, Mohsen Badiey, Thierry Bouwmans, Fragkiskos D. Malliaros
ICASSP, 2023
PDF / DOI
TL;DR Encoder-decoder GNN (TimeGNN) trained with a composite loss of MSE and a Sobolev graph smoothness operator, exploiting spatio-temporal correlations across graph topology for robust missing-entry recovery on real sensor and traffic datasets.
Recovery of Missing Sensor Data by Reconstructing Time-varying Graph Signals
Anindya Mondal, Mayukhmali Das, Aditi Chatterjee, Palaniandavar Venkateswaran
EUSIPCO, 2022
PDF / DOI / code
TL;DR Formulates missing wireless sensor data recovery as time-varying graph signal reconstruction minimising a Sobolev norm combining graph Laplacian smoothness across spatial and temporal dimensions, surpassing state-of-the-art by up to 54% at high missing-data rates.
Moving Object Detection for Event-based Vision using Graph Spectral Clustering
Anindya Mondal, Shashant R, Jhony H. Giraldo, Thierry Bouwmans, Ananda S. Chowdhury
ICCVW, 2021
PDF / DOI / code
TL;DR Maps asynchronous event streams to a k-NN graph, applies graph Laplacian spectral decomposition for eigenvector-based clustering (GSCEventMOD), and determines the optimal cluster count via spectral gap analysis, enabling fully unsupervised moving object detection for event cameras.

Education

University of Surrey, United Kingdom
Position : Doctor of Philosophy (Ph.D.) in Artificial Intelligence

Thesis: Unified Vision-Language Models, Generative AI, RL Post-training and Object Counting

Oct 2022 - Late 2026 (Expected)
Jadavpur University, Kolkata, India
Position : Bachelor of Engineering (B.E. Hons.) in Electronics & Telecom

Thesis: Graph Signal Processing and Event-based Vision

Aug 2018 - May 2022

Miscellanea

Academic Service

  • Reviewer: CVPR, ICCV, ECCV, NeurIPS, AAAI, SIGGRAPH, BMVC, IEEE TSP

Teaching & Mentorship

  • Supervision: 10+ MSc Dissertations (University of Surrey)
  • Lead Instructor: EEEM076 – Applied Machine Learning (University of Surrey)
  • Teaching Assistant:
    • EEEM071 – Advanced Topics in Computer Vision and Deep Learning (University of Surrey)
    • UKRI Centre for Doctoral Training in AI for Digital Media Inclusion

Template from Jon Barron.