Model techniques map
Techniquesmodel architecturemultimodal architecture

implementation detail · filed under model architecture

Reusing the image token for video frames

Video frames reuse the image token, with frame spans delimited by video start and end tokens.

Also called reusing image token for video frames.

source
1
model
1
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

the same image token is reused for video frames (frame spans are delimited by the video start/end tokens)

usedinference servingin GLM-5.3-FlashZ.ai

Filed alongside

Other methods under model architecture :: multimodal architecture.