Research, 2025
V-JEPA video representations
Adapting Meta's self-supervised V-JEPA features to video classification on a small GPU.
- Role
- Personal research project
- Type
- Research
- Links
- Source on GitHub
- benchmark
- UCF101
- attentive probing
- Probes
- constrained GPU memory
- Low VRAM
The problem
Self-supervised video models learn strong features without labels, but adapting them to a new task usually assumes plenty of GPU memory.
How I approached it
- 1
Probe, don't fine-tune
Attentive probes on top of frozen pretrained V-JEPA representations for UCF101 classification.
- 2
Experiment under limits
Compared temporal sampling strategies, probe configurations and batch sizes within constrained GPU memory.
What I built
- A PyTorch setup for training attentive probes on V-JEPA features within a constrained GPU memory budget.
The result
A practical comparison of what matters when adapting video foundation models on a budget.
Built with
- PyTorch
- V-JEPA
- CUDA
- Python