CIDR 2027 · Interactive artifact

One event query.
Three ways to execute it.

See how Concord substitutes a source-aligned transcript for video—or uses the transcript to restrict expensive video inference to a small fraction of the original lecture duration.

BaselineVideo → Localize
O1Transcript → Localize
O1 + O2 · Cross-modal rewrites

Lecture event localization

Find every interval where the lecturer creates a large, intact soap bubble that remains visible until it ruptures.

● recorded execution
VRA Loading recorded query…
    Now inspecting

    Loading…

    Source videoMIT 8.03 · Lecture 20
    129.8 s retained
    0:0082:25
    Processed video Reference events
    Source-aligned representation

    Transcript evidence

    Materialized View

    Candidate clip

    Reference Prediction

    Excerpt from MIT OpenCourseWare 8.03SC, Fall 2016. Candidate clip shown for research demonstration.

    Operator output

    Localized events

    Published evaluation · Table 2

    Less video, better localization

    All three queries use the same workload and output contract. Values below are frozen publication results, not recomputed in the browser.

    Video processedO2 lecture query
    Candidate recall4 of 4 events covered
    MLLM costrelative to video-only
    F1 scoreat tIoU ≥ .3
    Lecture event-localization results at the primary tIoU threshold.
    QueryVideoPrecisionRecallF1MLLM tokensCostTime
    Read the research

    Concord: A Video Relational Algebra for Cross-Modal Query Optimization

    Sultan Muratbek, Charisse Ivana Yeung, Chanwut Kittivorawong, and Alvin Cheung.

    MMDS repository ↗