Overview
Who this competition fits
Experienced ML researchers and engineers with strong backgrounds in multimodal systems, retrieval-augmented generation, and large language models. Ideal for those pushing the frontier of RAG with vision capabilities.
Read the original official blurb
Build multimodal RAG systems that can answer questions using text and images while minimizing hallucination. Part of the prestigious KDD Cup 2025.
Preparation
From registration to a first submission
- 01
Python
- 02
Large Language Models
- 03
Retrieval-Augmented Generation
- 04
Multimodal ML (vision + language)
- 05
Information retrieval fundamentals
- 06
Prompt engineering
Before you commit: Requires building end-to-end multimodal RAG pipelines that handle both text and image retrieval, cross-modal reasoning, and strict hallucination minimization. The KDD Cup competitive bar is extremely high.
Source
How this page was assembled
Competition information is structured from the official page. Scores are platform estimates for decision support; official rules take precedence.
- Official competition page
- AIcrowd
- Last checked
- Not recorded