SIGGRAPH Asia 2026 (ACM Transactions on Graphics)

ABACUS

Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation

Anindya Mondal*1,  Sauradip Nag*1,2,  Anjan Dutta1

1 University of Surrey, UK  •  2 Simon Fraser University, Canada

*Authors contributed equally

ABACUS overview: count-aware image generation (left) and count understanding across object, crowd, and referring expression counting (right)

Fig. 1. ABACUS overview. A single unified model performs count-aware image generation (left), object counting, crowd counting, and referring-expression counting (right) using text-only prompts with no benchmark-specific training.

Fast Forward Video Trailer

A quick walkthrough demonstrating ABACUS in action: bridging count understanding and count-faithful image generation.

Video 1. ABACUS Trailer. Demonstrating unified visual perception, adaptive zooming, objectness grounding, and cycle-consistent count generation.

One 3B Model. Four Tasks. Zero Benchmark-Specific Tuning.

We present ABACUS, a unified vision-language model that jointly addresses object counting, crowd counting, referring-expression counting, and count-faithful image generation within a single 3B-parameter model.

ABACUS introduces three complementary innovations:

  • Density-aware adaptive zooming paired with an objectness map from multi-head self-attention decomposition to spatially ground count predictions.
  • Boundary-aware count policy trained via GRPO with nested local, boundary, and global rewards to eliminate over- and undercounting at crop boundaries.
  • Cycle-consistent GRPO strategy in which the frozen understanding branch scores generated candidates on count-deviation and aesthetic quality, closing the understanding–generation synergy gap without any external critic or annotation.

ABACUS achieves state-of-the-art results across seven benchmarks spanning object counting (FSC-147, CARPK), crowd counting (ShanghaiTech A/B), referring-expression counting (REC-8K), count-faithful generation (CoCoCount, T2I-CompBench, GenEval), and count reasoning (CountQA), surpassing both task-specific specialists and larger generalist models.

Issues in Count Generation and Understanding: cardinality errors in diffusion models, coarse estimates in VLMs, and synergy gap in UMMs

Fig. 2. Issues in Count Generation and Understanding. (a) Text-to-image diffusion models lack any mechanism to verify output cardinality. (b) VLMs and MLLMs default to coarse magnitude estimates (e.g., “Greater than 100”) on dense scenes. (c) Existing UMMs support both tasks from a single model yet exhibit a synergy gap: the same model that correctly counts 4 apples cannot generate exactly 4.

Three Core Pillars of ABACUS

Bridging perception and generation through spatial grounding, boundary policy optimisation, and cycle-consistent reinforcement learning.

Density-aware adaptive zooming module
Pillar 1 • Understanding

Density-Aware Adaptive Zooming

A frozen GroundingDINO indicator $\phi(I)$ computes density score $s_d$. Dense scenes are recursively partitioned $2\times2$ to resolution $\gamma$; sparse scenes are evaluated in a single pass. A boundary-aware GRPO policy with nested local, boundary, and global rewards eliminates double-counting at tile borders.

Infusing objectness in MLLM via MHSA decomposition
Pillar 2 • Spatial Grounding

Objectness Map via MHSA

Exploiting linear decomposability of Multi-Head Self-Attention in the language model. Per-head isolation and learned affine alignment produce spatial distributions over visual token positions, supervised by $\mathcal{L}_{\mathrm{obj}}$ with Gaussian targets to internalise instance structure without bounding-box labels.

Count-aware image enhancement via cycle-consistent GRPO
Pillar 3 • Generation

Cycle-Consistent GRPO

The generation branch samples $N$ candidates from a count prompt. The frozen understanding branch self-critiques candidates on count deviation and aesthetic quality. The composite reward updates the generation LoRA via GRPO advantage estimation, closing the synergy gap with zero external supervision.

ABACUS unified model architecture pipeline

Fig. 3. ABACUS unified model architecture. Unified VLM foundation coupling an InternViT visual encoder, Qwen2 language backbone, multimodal connector, and SANA diffusion transformer with DC-AE pixel decoder. A single parameter-efficient adapter bridges understanding and generation.

Comprehensive Benchmark Evaluations

Outperforming task-specific specialists and generalist foundation models across 7 standard benchmarks with a single unified 3B checkpoint.

Method Sup. FSC-147 Val FSC-147 Test CARPK Test ShanghaiTech A ShanghaiTech B CountQA
MAE ↓RMSE ↓ MAE ↓RMSE ↓ MAE ↓RMSE ↓ MAE ↓RMSE ↓ MAE ↓RMSE ↓ EM % ↑
Specialist Counting Models
CountGD++P 12.1447.518.3927.03——116.0234.028.050.0—
T2ICountP 13.7858.7811.7697.868.6113.47—————
CountSE†P ——7.8482.99——129.7258.3———
CAD-GDP ——10.3586.88———————
VLM-based Counting Models
GPT-5.5Z 25.8779.3425.17162.024.3336.50215.67412.8958.92102.3425.03
Show-oZ 37.87105.5546.26129.5342.1563.22312.45587.3389.34156.787.85
Janus Pro 7BZ 43.56110.2335.7099.9634.5251.78278.91523.4476.23134.566.98
UniLIP-3BZ 30.19103.0726.44103.9826.8733.48243.15424.3363.7997.049.23
WS-COC-7BI 14.7754.2413.9197.2810.3915.83128.9232.934.257.08.44
ABACUS-3B (Ours)Z 5.7126.46 5.0327.03 8.4110.84 78.59139.88 14.7525.08 15.30

Table 1. Comprehensive evaluation of Object Counting across FSC-147, CARPK, ShanghaiTech (SHT A & B), and Count Reasoning on CountQA. † indicates few-shot visual exemplar methods. Supervision: Point-level (P), Image-level (I), Zero-shot/MLLM (Z).

Method CoCoCount T2I-CompBench GenEval
YOLOv9 ↑Human ↑ Human ↑ YOLOv9 ↑Human ↑Aesthetic ↑
Specialist Count Generation
CountGen505448464445
BoundedAttn293035211810
Counting Guidance21222216117
VLM-based Count Generation
BAGEL364132443943
Janus Pro-7B273225303358
UniLIP-3B343930364061
ABACUS-3B (Ours) 7177 65 949589

Table 2. Count Generation evaluation across CoCoCount, T2I-CompBench, and GenEval. Exact-match accuracy (%) with YOLOv9 detector and human annotator study, alongside Aesthetic Quality.

Method Backbone Fine-tuned on REC MAE ↓ RMSE ↓
Specialist Counting Models (Detection-based with Box Supervision)
ZSCSwin-T13.0029.07
TFOCViT-B—17.2732.68
CounTXViT-B/1611.8425.62
GroundingDINO (FT)Swin-T8.8821.95
GrRECSwin-T6.5019.79
Unified Multimodal Foundation Models
UniLIP-3BUniLIP-3B—13.7525.91
ABACUS-3B (Ours)UniLIP-3B— 7.6715.84

Table 3. Referring Expression Counting on REC-8K test set ($n=3{,}153$ pairs). Evaluated text-only without bounding-box supervision or benchmark fine-tuning. Surpasses fine-tuned GDINO and achieves the lowest RMSE (15.84) across all methods.

Visual Comparison Across Density Regimes

Comparing ABACUS against leading specialists and generalist VLMs on counting and generation.

In-Depth Empirical Analysis

Investigating training convergence, density inference latency, cardinality limits, and failure boundaries.

Training Strategy on CoCoCount (Fig. 6)
Ablation of Training Strategy on CoCoCount

Cycle-consistent GRPO converges to 71% exact-match on CoCoCount, outperforming open-loop GRPO with an external counter (62%) and LoRA SFT alone (45%), demonstrating that self-reinforcing co-adaptation between generator and understanding critic is essential.

Inference Cost & Memory in Dense Scenes (Fig. 7)
Inference latency and memory overhead across density regimes

Adaptive zooming preserves single-pass latency (310 ms, 7.2 GB) on sparse scenes ($<20$ objects), with latency scaling gracefully to $1.2\times$ (380 ms) for 20–100 objects, $2.0\times$ (620 ms) for 100–500 objects, and $3.7\times$ (1150 ms) for extreme crowds ($>500$).

Cardinality Scaling & Physical Packing Limits
Requested CountExact AccuracyAesthetics (0–1)
10 objects72.4%0.91
50 objects71.1%0.72
100 objects68.3%0.65

Evaluating generative fidelity up to 100 instances. While physical packing constraints in diffusion models cause slight degradation at extreme counts, ABACUS maintains over 68% accuracy without sacrificing composition realism.

Failure Modes & Boundaries (Fig. 8)
Failure cases of ABACUS on low resolution and microscopy cell counting

Under severe resolution degradation ($<224$ px), InternViT’s $14\times14$ patch grid becomes too coarse to separate individual objects. Specialized domain shifts such as microscopic cell counting also require domain-specific LoRA adaptation.

Citation

If you find ABACUS useful in your research, please cite our camera-ready paper:

@article{mondal2026abacus,
  title   = {ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation},
  author  = {Mondal, Anindya and Nag, Sauradip and Dutta, Anjan},
  journal = {ACM Transactions on Graphics (TOG)},
  year    = {2026}
}

Acknowledgements

The authors gratefully acknowledge NVIDIA Corporation for support through the NVIDIA Academic Grant Program, which provided computational resources for this research. The authors also acknowledge the use of resources provided by the Isambard-AI National AI Research Resource (AIRR), operated by the University of Bristol and funded by the UK Government’s Department for Science, Innovation and Technology (DSIT) through UK Research and Innovation and the Science and Technology Facilities Council [ST/AIRR/I-A-I/1023].