The multimodal AI that understands your video library

Our AI systems watch and listen to your footage the way a person would, then put that understanding to work. Developed in-house, with industry-leading security and rightsholder protections.

Perception is only the start. What our systems see and hear feeds context, and context feeds the intelligence that answers questions, cuts clips and organizes your library.

Our AI systems at a glance

One library, two engines

Vulpus finds the moment. Cetus explains and enriches it. Every video you index flows through both, and you can use either on its own.

Engine 1 · Vulpus

Search without tags

Contextual understanding across visuals, audio and dialogue. Describe a moment in natural language and Vulpus retrieves the exact shot, with no labels or metadata needed.

Engine 2 · Cetus

Metadata that explains itself

Scene-level descriptions with confidence scores, enriched at asset and scene level and mapped to your own taxonomy. Built for teams whose workflows run on metadata.

Everything our AI systems identify

Live on Imaginario AI Person IDCelebritiesFacial expressionsHuman emotionsActionsLocationsObjectsBrandsSFXSpeech · 100+ languagesScene context & descriptionsChaptersSummariesKeywordsAuto clippingShot detectionAuto trackingTranscriptions & captionsSynopsis New · How it was made, moods and narrative Audiovisual languageMoodsNarrative structure & beat analysisCompliance flagsOverlaysTitle cardsLogo bugsBlack screens

For Enterprise customers, we can create custom models based on your footage and map outputs to your own vocabulary or editorial taxonomy.

Vulpus

Our pioneering multimodal AI system, designed to analyze videos through visuals, audio, and dialogue while understanding the passing of time.

By processing and integrating data from each of these modalities and utilizing an advanced search and discovery engine, Vulpus offers exceptional search capabilities across extensive video libraries without the need for labels or metadata.

Vulpus sets a new standard for video content retrieval and analysis, making it easier and more efficient to index, curate and manage libraries of any size.

Key info

ModalitiesVisual, audio, dialogue
AvailabilityImaginario app, API

Capabilities

  • Scalable processing and analysis of all video content (any length / any file size)
  • Selective analysis by modality (e.g. analyze only visuals or only audio)
  • Customizable clip retrieval (choose number of results, length of clips)
  • Near-instant result retrieval after initial indexing
  • Natural language multimodal search (combine object and audio search, e.g. “a busy restaurant with piano music playing”)
  • Customizable face search

Selected use cases

Cetus

Our second multimodal AI system, specifically developed to increase explainability and provide a modern complement for video labels and metadata.

Utilizing the analyzes of Vulpus, Cetus provides descriptions of video content at a scene level, taking into account the action, characters, speech and context.

Key info

ModalitiesVisual, audio, dialogue, faces, temporal understanding, fine-tuning for private taxonomies and schemas
AvailabilityImaginario app, API

Capabilities

  • Scalable processing and analysis of all video content (any length / any file size)
  • Natural-language descriptions of any scene, taking into account visuals, audio and dialogue
  • Instant description generation
  • Percentage confidence score based on search term

See it in action

Frequently asked questions

Is there a limit to how much I can search my library?

Our packages are based on the minutes of video content you import every month, not the number of searches. Once you have added your video content to our system, you are free to search as many times as you like, across all modalities.

Our AI system understands your videos in the same way a human would, so you can search for anything that exists within your footage.

This could be something specific (such as an actor or a model of car), or something more thematic such as “winter” or “suspense”.

Our AI also understands the relationship between people, objects, places, emotions, and actions. So if you want to locate a specific shot you can combine elements, for example “a group of people running on the beach”.

We add intelligence and context over and above the previous generation of video search services, so you can add details and relational data to your searches to find exactly what you need.

For example, you’re preparing some social clips for Valentine’s Day so you want to find a specific scene where two characters kiss in a crowded street. With other visual search systems, you would be able to search for “man”, “woman”, “kiss” and maybe “crowd”. With our system, you can search for “A man and a woman kissing in a crowded street” and you will find exactly the right clip.

Intelligence is what connects understanding your library with acting on it. Imaginario indexes visuals, speech, sound and people at shot level, and the platform acts on that understanding: cuts and new edits, curated collections, insights from content performance, compliance reviews and much more. It is intelligence on top of your video libraries and projects, not just a chat window. You can ask for something specific, such as a person or one line from one interview, or something thematic, such as the strongest customer quotes from this year’s webinars. Because the index understands the relationship between people, objects, places, emotions and actions, the answers come back grounded in what actually happens on screen, ready to become an edit, a report or a decision.
No. General assistants accept documents and images, not video files, and they cannot index a private library. Ask your library gives you the same conversational interface on footage Imaginario has actually indexed: visuals, speech, sound and people at shot level. That is the difference between chat and intelligence: answers you can act on.
Twelve Labs provides strong video understanding models through an API, and that kind of technology is one part of our stack. Imaginario takes a 360-degree approach on top of it: multimodal understanding plus shot-level and scene-level descriptions, with the generative cost included in indexing rather than billed as separate infrastructure. Imaginario is also a full video intelligence system rather than an API alone. Developers get the API, and non-technical teams get the full app for search, curation and clipping. You can see the full comparison on our Twelve Labs page.

Get started today

Enjoy thousands of minutes of AI analysis, transcriptions, search, and auto-captions, plus much more.