Skip to content

Research, 2025

V-JEPA video representations

Adapting Meta's self-supervised V-JEPA features to video classification on a small GPU.

Role
Personal research project
Type
Research
benchmark
UCF101
attentive probing
Probes
constrained GPU memory
Low VRAM

The problem

Self-supervised video models learn strong features without labels, but adapting them to a new task usually assumes plenty of GPU memory.

How I approached it

  1. 1

    Probe, don't fine-tune

    Attentive probes on top of frozen pretrained V-JEPA representations for UCF101 classification.

  2. 2

    Experiment under limits

    Compared temporal sampling strategies, probe configurations and batch sizes within constrained GPU memory.

What I built

  • A PyTorch setup for training attentive probes on V-JEPA features within a constrained GPU memory budget.

The result

A practical comparison of what matters when adapting video foundation models on a budget.

Built with

  • PyTorch
  • V-JEPA
  • CUDA
  • Python