Jag: Direct Box Prediction for Efficient Visual Grounding

Adapt multimodal representations to geometry.
Predict the complete bounding box in one forward pass.

Xiuyuan Zhu1,2Ke Lu1,3Hao Wu1,2Zijin Du1Dongming Zhang2Jian Xue1,*
1 University of Chinese Academy of Sciences2 State Key Laboratory of Communication Content Cognition3 Peng Cheng Laboratory
0.8BMultimodal backbone
65.97 msMean prediction latency
3.62×Faster than LocateAnything
2.52 GiBPeak GPU process memory

Latency and memory are measured in the single-request comparison.

01

Abstract

Visual grounding requires an image, a referring expression, and a bounding box. Autoregressive multimodal models often express that box as a sequence of coordinate tokens, adding decoding steps to the image–language computation. Jag learns to predict the box directly from the model’s existing input representations. A lightweight regression head reads the last valid input-token state and returns all four continuous box values together. Geometric losses adapt both the multimodal representations and their mapping to coordinates. The resulting model improves over its base model across all five evaluated test splits, with competitive accuracy and lower single-request latency than the compared specialized grounding models.

02

Method

Jag keeps the pretrained image–language computation and adapts it for continuous localization, without generating coordinate text.

Jag replaces sequential coordinate decoding with direct continuous box prediction from the last valid input-token state. L1 and GIoU losses train the regression head and adapt the multimodal representations, improving grounding accuracy and inference efficiency.
Jag adapts multimodal representations for direct box prediction. Geometric supervision turns an existing input state into a complete bounding box in one forward pass, combining accurate grounding with efficient inference. View full size ↗
A

Read the input state

The readout follows the full image and referring expression, so the box prediction reuses their multimodal representation.

B

Supervise the geometry

L1 and GIoU losses train continuous box centers and dimensions, connecting the output directly to coordinate error and region overlap.

C

Adapt the representations

Head adaptation precedes joint fine-tuning of the language backbone, visual merger, and regression head. The remaining vision encoder stays frozen.

03

Grounding Accuracy

Jag improves over Base and NExT-Chat on every evaluated split and achieves the highest accuracy on three of the five splits in this comparison.

Referring expression comprehensionAcc@0.5 (%) ↑
ModelRefCOCORefCOCO+RefCOCOg
testAtestBtestAtestBtest
Base Qwen3.5-0.8B84.2774.7276.5362.5777.96
NExT-Chat89.6677.0483.7666.1979.28
LocateAnything93.2389.2688.0079.5788.54
Jag Ours93.9088.3690.6980.0887.72

A prediction is correct when its intersection over union with the target box is at least 0.5. Bold marks the best value in each column.

Consistent gains over Base. The improvement extends to every test split, including RefCOCO+, where expressions restrict location words and emphasize visual appearance.

Competitive with specialist models. Jag exceeds NExT-Chat on all five splits and LocateAnything on three, while LocateAnything leads on RefCOCO testB and RefCOCOg.

04

Inference Efficiency

Direct regression returns all four coordinates after one multimodal forward pass. No subsequent coordinate generation is required.

Latency versus peak GPU memory: Jag is the lowest-latency model at 65.97 milliseconds and uses 2.52 GiB, less memory than NExT-Chat, LocateAnything, and Hi-Token. Base uses slightly less memory than Jag.
Jag achieves the lowest measured latency. The shaded region contains combinations with both higher latency and higher memory use than Jag.
Single-request inferenceBatch size 1
ModelLatency ms ↓Memory GiB ↓Jag speedup
Base1,293.012.2519.60×
NExT-Chat120.0115.991.82×
LocateAnything239.059.663.62×
Hi-Token616.907.919.35×
Jag65.972.521.00×

Mean end-to-end prediction time and peak GPU process memory. Speedup compares each model’s latency with Jag. Model loading and warmup are excluded.

3.62× faster with 73.88% less GPU memory than LocateAnything in this comparison.

05

Ablation Studies

Jointly adapting the language backbone, visual merger, and regression head improves localization over fitting the head alone. Alternative readouts and training schedules remain competitive.

Component ablations comparing joint adaptation, head-only training, readout positions, and training schedules using Acc@0.5.
Component ablations examine the effect of representation adaptation, the readout, and the training schedule. Dashed lines mark the autoregressive reference.
06

Qualitative Results

Examples from RefCOCOg compare Jag and Base against the annotated region, including shared successes and cases where either model fails.

Four RefCOCOg examples: both models correctly localize a table; only Jag correctly localizes an elephant with its trunk up; only Base correctly localizes an airplane; both models fail on an expanse of table above a bowl.
Green: annotated box. Blue: Jag. Dashed gray: Base. Each expression refers to one target region.
08

Citation

If you use Jag, you can cite the project:

@misc{zhu2026jag,
  title  = {Jag: Direct Box Prediction for Efficient Visual Grounding},
  author = {Zhu, Xiuyuan and Lu, Ke and Wu, Hao and Du, Zijin
            and Zhang, Dongming and Xue, Jian},
  year   = {2026},
  url    = {https://github.com/xyzzzh/Jag}
}
09

Contact

For questions about the paper, code, or model, please contact us by email.