Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation
1 University of Surrey, UK • 2 Simon Fraser University, Canada
*Authors contributed equally
Fig. 1. ABACUS overview. A single unified model performs count-aware image generation (left), object counting, crowd counting, and referring-expression counting (right) using text-only prompts with no benchmark-specific training.
A quick walkthrough demonstrating ABACUS in action: bridging count understanding and count-faithful image generation.
Video 1. ABACUS Trailer. Demonstrating unified visual perception, adaptive zooming, objectness grounding, and cycle-consistent count generation.
We present ABACUS, a unified vision-language model that jointly addresses object counting, crowd counting, referring-expression counting, and count-faithful image generation within a single 3B-parameter model.
ABACUS introduces three complementary innovations:
ABACUS achieves state-of-the-art results across seven benchmarks spanning object counting (FSC-147, CARPK), crowd counting (ShanghaiTech A/B), referring-expression counting (REC-8K), count-faithful generation (CoCoCount, T2I-CompBench, GenEval), and count reasoning (CountQA), surpassing both task-specific specialists and larger generalist models.
Fig. 2. Issues in Count Generation and Understanding. (a) Text-to-image diffusion models lack any mechanism to verify output cardinality. (b) VLMs and MLLMs default to coarse magnitude estimates (e.g., “Greater than 100”) on dense scenes. (c) Existing UMMs support both tasks from a single model yet exhibit a synergy gap: the same model that correctly counts 4 apples cannot generate exactly 4.
Bridging perception and generation through spatial grounding, boundary policy optimisation, and cycle-consistent reinforcement learning.
A frozen GroundingDINO indicator $\phi(I)$ computes density score $s_d$. Dense scenes are recursively partitioned $2\times2$ to resolution $\gamma$; sparse scenes are evaluated in a single pass. A boundary-aware GRPO policy with nested local, boundary, and global rewards eliminates double-counting at tile borders.
Exploiting linear decomposability of Multi-Head Self-Attention in the language model. Per-head isolation and learned affine alignment produce spatial distributions over visual token positions, supervised by $\mathcal{L}_{\mathrm{obj}}$ with Gaussian targets to internalise instance structure without bounding-box labels.
The generation branch samples $N$ candidates from a count prompt. The frozen understanding branch self-critiques candidates on count deviation and aesthetic quality. The composite reward updates the generation LoRA via GRPO advantage estimation, closing the synergy gap with zero external supervision.
Fig. 3. ABACUS unified model architecture. Unified VLM foundation coupling an InternViT visual encoder, Qwen2 language backbone, multimodal connector, and SANA diffusion transformer with DC-AE pixel decoder. A single parameter-efficient adapter bridges understanding and generation.
Outperforming task-specific specialists and generalist foundation models across 7 standard benchmarks with a single unified 3B checkpoint.
| Method | Sup. | FSC-147 Val | FSC-147 Test | CARPK Test | ShanghaiTech A | ShanghaiTech B | CountQA | |||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MAE ↓ | RMSE ↓ | MAE ↓ | RMSE ↓ | MAE ↓ | RMSE ↓ | MAE ↓ | RMSE ↓ | MAE ↓ | RMSE ↓ | EM % ↑ | ||
| Specialist Counting Models | ||||||||||||
| CountGD++ | P | 12.14 | 47.51 | 8.39 | 27.03 | — | — | 116.0 | 234.0 | 28.0 | 50.0 | — |
| T2ICount | P | 13.78 | 58.78 | 11.76 | 97.86 | 8.61 | 13.47 | — | — | — | — | — |
| CountSE† | P | — | — | 7.84 | 82.99 | — | — | 129.7 | 258.3 | — | — | — |
| CAD-GD | P | — | — | 10.35 | 86.88 | — | — | — | — | — | — | — |
| VLM-based Counting Models | ||||||||||||
| GPT-5.5 | Z | 25.87 | 79.34 | 25.17 | 162.0 | 24.33 | 36.50 | 215.67 | 412.89 | 58.92 | 102.34 | 25.03 |
| Show-o | Z | 37.87 | 105.55 | 46.26 | 129.53 | 42.15 | 63.22 | 312.45 | 587.33 | 89.34 | 156.78 | 7.85 |
| Janus Pro 7B | Z | 43.56 | 110.23 | 35.70 | 99.96 | 34.52 | 51.78 | 278.91 | 523.44 | 76.23 | 134.56 | 6.98 |
| UniLIP-3B | Z | 30.19 | 103.07 | 26.44 | 103.98 | 26.87 | 33.48 | 243.15 | 424.33 | 63.79 | 97.04 | 9.23 |
| WS-COC-7B | I | 14.77 | 54.24 | 13.91 | 97.28 | 10.39 | 15.83 | 128.9 | 232.9 | 34.2 | 57.0 | 8.44 |
| ABACUS-3B (Ours) | Z | 5.71 | 26.46 | 5.03 | 27.03 | 8.41 | 10.84 | 78.59 | 139.88 | 14.75 | 25.08 | 15.30 |
Table 1. Comprehensive evaluation of Object Counting across FSC-147, CARPK, ShanghaiTech (SHT A & B), and Count Reasoning on CountQA. † indicates few-shot visual exemplar methods. Supervision: Point-level (P), Image-level (I), Zero-shot/MLLM (Z).
| Method | CoCoCount | T2I-CompBench | GenEval | |||
|---|---|---|---|---|---|---|
| YOLOv9 ↑ | Human ↑ | Human ↑ | YOLOv9 ↑ | Human ↑ | Aesthetic ↑ | |
| Specialist Count Generation | ||||||
| CountGen | 50 | 54 | 48 | 46 | 44 | 45 |
| BoundedAttn | 29 | 30 | 35 | 21 | 18 | 10 |
| Counting Guidance | 21 | 22 | 22 | 16 | 11 | 7 |
| VLM-based Count Generation | ||||||
| BAGEL | 36 | 41 | 32 | 44 | 39 | 43 |
| Janus Pro-7B | 27 | 32 | 25 | 30 | 33 | 58 |
| UniLIP-3B | 34 | 39 | 30 | 36 | 40 | 61 |
| ABACUS-3B (Ours) | 71 | 77 | 65 | 94 | 95 | 89 |
Table 2. Count Generation evaluation across CoCoCount, T2I-CompBench, and GenEval. Exact-match accuracy (%) with YOLOv9 detector and human annotator study, alongside Aesthetic Quality.
| Method | Backbone | Fine-tuned on REC | MAE ↓ | RMSE ↓ |
|---|---|---|---|---|
| Specialist Counting Models (Detection-based with Box Supervision) | ||||
| ZSC | Swin-T | 13.00 | 29.07 | |
| TFOC | ViT-B | — | 17.27 | 32.68 |
| CounTX | ViT-B/16 | 11.84 | 25.62 | |
| GroundingDINO (FT) | Swin-T | 8.88 | 21.95 | |
| GrREC | Swin-T | 6.50 | 19.79 | |
| Unified Multimodal Foundation Models | ||||
| UniLIP-3B | UniLIP-3B | — | 13.75 | 25.91 |
| ABACUS-3B (Ours) | UniLIP-3B | — | 7.67 | 15.84 |
Table 3. Referring Expression Counting on REC-8K test set ($n=3{,}153$ pairs). Evaluated text-only without bounding-box supervision or benchmark fine-tuning. Surpasses fine-tuned GDINO and achieves the lowest RMSE (15.84) across all methods.
Comparing ABACUS against leading specialists and generalist VLMs on counting and generation.
Fig. 4. Qualitative comparison on count understanding. ABACUS (Ours) tracks ground truth across all density regimes, from sparse (GT: 4) to extremely dense (GT: 1400). CountGD++ overcounts in dense scenes (900 for GT: 298); T2I Count catastrophically fails on out-of-distribution layouts (0 for GT: 1231); WS-COC and UniLIP-3B default to coarse magnitude estimates.
Fig. 5. Qualitative comparison on count generation. ABACUS achieves exact or near-exact counts while preserving natural spatial arrangement. UniLIP-3B systematically undercounts; BAGEL overcounts with unnatural compositions; CountGen produces rigid grid patterns; Counting Guidance exhibits mode collapse (39 donuts for prompt 21, 364 balls for prompt 28).
Investigating training convergence, density inference latency, cardinality limits, and failure boundaries.
Cycle-consistent GRPO converges to 71% exact-match on CoCoCount, outperforming open-loop GRPO with an external counter (62%) and LoRA SFT alone (45%), demonstrating that self-reinforcing co-adaptation between generator and understanding critic is essential.
Adaptive zooming preserves single-pass latency (310 ms, 7.2 GB) on sparse scenes ($<20$ objects), with latency scaling gracefully to $1.2\times$ (380 ms) for 20–100 objects, $2.0\times$ (620 ms) for 100–500 objects, and $3.7\times$ (1150 ms) for extreme crowds ($>500$).
| Requested Count | Exact Accuracy | Aesthetics (0–1) |
|---|---|---|
| 10 objects | 72.4% | 0.91 |
| 50 objects | 71.1% | 0.72 |
| 100 objects | 68.3% | 0.65 |
Evaluating generative fidelity up to 100 instances. While physical packing constraints in diffusion models cause slight degradation at extreme counts, ABACUS maintains over 68% accuracy without sacrificing composition realism.
Under severe resolution degradation ($<224$ px), InternViT’s $14\times14$ patch grid becomes too coarse to separate individual objects. Specialized domain shifts such as microscopic cell counting also require domain-specific LoRA adaptation.
If you find ABACUS useful in your research, please cite our camera-ready paper:
@article{mondal2026abacus,
title = {ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation},
author = {Mondal, Anindya and Nag, Sauradip and Dutta, Anjan},
journal = {ACM Transactions on Graphics (TOG)},
year = {2026}
}
The authors gratefully acknowledge NVIDIA Corporation for support through the NVIDIA Academic Grant Program, which provided computational resources for this research. The authors also acknowledge the use of resources provided by the Isambard-AI National AI Research Resource (AIRR), operated by the University of Bristol and funded by the UK Government’s Department for Science, Innovation and Technology (DSIT) through UK Research and Innovation and the Science and Technology Facilities Council [ST/AIRR/I-A-I/1023].