NeurIPS 2026

LoHi

Rethinking Long-Video Efficiency: A Joint Allocation Perspective on Frames, Pixels, and Front-End Latency

Trade pixels for frames. Decode half as many. Answer faster and better.

Sixun Dong1  ·  Wei Li1  ·  Andong Deng1  ·  Qi Qian2  ·  Victor Zhu3  ·  Zhengping Ji3  ·  Chen Chen1✉

1University of Central Florida   2Meta Reality Labs   3Axon  ·  ✉ corresponding author

Paper Blog arXiv soon Code soon

We are building longvideo-eval, an open-source harness that evaluates long-video VLMs and efficiency methods on one accuracy–cost axis, with decoding, vision encoding and prefill all measured. LoHi will be released as part of it.

Prior work: keep native resolution, decide what to throw away
Keyframe selectiondecode 256 frames, keep 16
256 decoded6.0 s to first token60.8% VideoMME
Token pruningencode 256 native frames, drop 94% of tokens
256 decoded9.0 s to first token59.3% VideoMME
Our recipe: lower the resolution, keep the whole timeline
Low-Res-Base256 frames at ¼ resolution, no selection at all
256 decoded6.0 s to first token64.4% VideoMME
LoHi128 frames at ¼ resolution, plus 8 of them at native
128 decoded3.9 s to first token66.8% VideoMME
native-resolution framenative frame, mostly prunedlow-resolution framehigh-resolution pickdecoded, then discardedbar height = resolution · one 60-min video · ~5,760 visual tokens · Qwen3-VL-4B
Three lessons

What actually limits long-video VLMs

Efficient long-video VLMs usually ask which tokens to keep at native resolution. We ask how to spend a fixed budget across frame count, per-frame resolution, and front-end decoding. Three lessons answer it. Every VLM spends a fixed visual-token budget B on one configuration (N, r): N frames at per-frame resolution scale r, with N · τ(r) ≤ B. Before the tokens ever reach the model, a video front-end must decode those frames from the compressed file. Play with each lesson below.

L1

Dense low-resolution sampling is the new recipe

At the same token budget, 256 frames at ¼ resolution beat 16 frames at native resolution, and beat dedicated keyframe-selection and token-pruning methods too. We call this Low-Res-Base.

One 60-min video, one ~5,760-token budget, five ways to spend it bar height = per-frame resolutionlow-res framenative-res frameLoHi Hi-I pickdecoded, then discarded
00:00timeline →60:00
VideoMME accuracy57.78%
time to first token—

decoded 16sent to VLM 16 @ 1.0×tokens ≈5,760avg. of 3 benchmarks 51.76

Numbers: Qwen3-VL-4B, VideoMME (w/o subtitles). TTFT is measured on an H100 and includes CPU decoding; the vanilla 16-frame row has no TTFT entry in the paper. Keyframe selection shows BOLT, the strongest keyframe method; token pruning shows FlashVid.

Takeaway (L1). Under a strict token budget, trading per-frame resolution for dense temporal coverage consistently outperforms the default sparse, native-resolution recipe. Dense, low-resolution sampling should serve as the new foundational baseline for long-video VLMs.
L2

Resolution sensitivity is task-dependent

Shrinking a frame is surprisingly cheap: people, actions and the layout of a scene survive long after the pixels get blocky, and the VLM keeps up. Fine text and small attributes do not. Turn the dial and watch what disappears first.

The resolution dial: same frame, same token budget tokens per frame ∝ r²frames you can afford = budget ÷ tokens per frame
1.00.50.250.1250.0625
close-up of the sign, shown at native pixel size
tokens / frame360
frames in ~5,760 tokens16

    From 1.0 to 0.25, the paper measures
    −9.9 Attribute Perception · −10.1 OCR
    while Action Recognition and Temporal Reasoning stay robust (Fig. 1b).
    At the same budget, the winner flips
    [email protected] wins R 72.73 vs 71.89 · [email protected] wins ¬R 61.76 vs 58.54
    R = resolution-sensitive questions (OCR, Attribute). LoHi-SemDiv beats both: 77.06 on R, 63.12 on ¬R.

    The frame is a synthetic illustration rendered in your browser; accuracy numbers are from the paper (Qwen3-VL-4B, VideoMME). Scales below 0.25 are shown only to build intuition and were not evaluated.

    Takeaway (L2). Under a fixed token budget, no single (N, r) optimally serves all task types; any globally fixed allocation is suboptimal on at least one subset.
    L3

    Front-end latency cannot be overlooked in long videos

    A video file is not a stack of pictures. Codecs store a full I-frame at the start of each Group of Pictures (GoP), and every following P-frame only stores what changed since the previous frame. To show any P-frame, the decoder must start at its I-frame and replay every frame in between.

    Click any frame: how much must be decoded to get it? I-frame (self-contained)P-frame (stores changes)had to decodeframe you wanted

    Measured decode latency (CPU decord)

    Dense sampling touches every GoP and grows with video length. Uniform sampling of a fixed N caps both.

    Why it matters

    Decode-then-select keyframe methods (AKS, BOLT, CLIP top-K) first build a candidate pool by densely decoding the video, then score it. On a one-hour clip the decoder alone runs for tens of seconds while the GPU sits idle. Token pruning decodes 256 native-resolution frames and runs the full vision encoder before discarding 93.75% of tokens.

    LoHi bounds the pool: it decodes exactly N = 128 frames once, regardless of length, and reuses them at two resolutions. LoHi-Anchor even reads I-frame positions as a free scene-change hint, without decoding extra pixels.

    Takeaway (L3). Under dense candidate-pool decoding such as FPS=1, front-end latency grows with video length and eventually dominates end-to-end runtime. Long-video efficiency methods must therefore bound the size of the decoded candidate pool, not just the visual-token budget passed to the VLM.
    Method

    LoHi: low-resolution video + high-resolution images

    LoHi splits one configuration (N, r) into two complementary streams that share one budget and one decoding pass, and routes them through the VLM's own video and image pathways. Training-free, single forward pass, no architectural change.

    Overview of the LoHi framework compared with previous methods
    Overview of LoHi. Compared with previous work (A), which prunes or selects within a native-resolution pathway, LoHi (B) decomposes one decoded video into two complementary streams: dense low-resolution Lo-V via the video pathway and sparse high-resolution Hi-I via the image pathway. (C) shows the LoHi-SemDiv selector, which picks Hi-I indices via greedy MAP.

    Plug-and-play Hi-I selectors

    LoHi-Uniform

    selection cost: zero

    Picks the K Hi-I frames at regular intervals directly from the Lo-V grid. Pure indexing, and the fundamental baseline of the framework.

    LoHi-Anchor

    selection cost: free codec metadata

    Encoders start GoPs with I-frames, which can align with scene changes or strong motion. For each uniform position, Anchor finds the nearest I-frame from pre-computed indices and selects the Lo-V grid frame closest to it, with no additional pixel decoding.

    LoHi-SemDiv

    selection cost: one CLIP pass on 128 small frames

    A quality–similarity DPP over CLIP embeddings: L = diag(q) EE⊤ diag(q), with q the sharpened query relevance. Greedy maximization of log det(LS + I) picks frames that are both query-relevant and mutually diverse, with a (1 − 1/e) guarantee.

    Results

    The most accurate, and the fastest to answer

    Qwen3-VL-4B, every method at the token budget of the default 16-frame recipe (~5,760 tokens). Keyframe selection and token pruning decode 256 native-resolution frames; LoHi decodes 128.

    Accuracy vs. time to first token

    VideoMME, Qwen3-VL-4B, matched ~5,760-token budget. Up and to the left is better. Hover a point for details.
    LoHi (ours) Keyframe selection Token pruning Low-Res-Base (reference)
    MethodDecoded framesVideoMMEMLVULVBenchAvg.
    Qwen3-VL-4B, [email protected]1657.7858.8938.6151.76
    Low-Res-Base, [email protected]25664.4469.3643.7159.17
    Keyframe selection (best: BOLT)256 → 1660.8168.1142.4857.13
    Token pruning (best: FlashVid)25659.3063.7034.0952.36
    LoHi-Uniform12865.7867.6244.0959.16
    LoHi-Anchor12865.8968.4044.4259.57
    LoHi-SemDiv12866.8174.3845.9062.36

    LoHi-SemDiv improves the average over the default recipe by +10.6 and over the strongest prior method by +5.2. The blog has the full tables: every baseline, latency breakdown, token pruning vs. resize, ablations and five more backbones.

    A detail on the interviewee's chin that only the high-resolution frames recover
    Blog

    The full story behind LoHi

    Why the video decoder is the hidden bottleneck, why a plain resize beats token pruning on long videos, what happens when the selector misses, and all the tables that did not fit here.

    Read the blog post
    Citation

    BibTeX

    @inproceedings{dong2026lohi,
      title     = {Rethinking Long-Video Efficiency: A Joint Allocation Perspective on Frames, Pixels, and Front-End Latency},
      author    = {Dong, Sixun and Li, Wei and Deng, Andong and Qian, Qi and Zhu, Victor and Ji, Zhengping and Chen, Chen},
      booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
      year      = {2026},
      url       = {https://openreview.net/forum?id=a9xLyT4hG4}
    }

    Questions or collaboration: [email protected]