An editor remembers the shot clearly. A traveller stands alone on a concourse as the light moves across the glass. It appeared briefly in an early assembly, perhaps from the second shoot day. The editor does not remember the clip name, the camera roll or where the material was eventually filed.
The search begins in familiar ways. Open the old project. Scrub through an earlier cut. Check the treatment for the original visual reference. Browse folders of proxies and scan rows of thumbnails. Ask the assistant editor whether they remember the take. If the shot came from another campaign, search the wider archive and hope that the folder structure provides a clue.
Nothing about this suggests a badly run production. The media may have been ingested correctly, logged consistently and stored exactly where it belongs. The difficulty is simply that the editor remembers the meaning of the shot, while the archive remembers its filename.
Filenames describe storage, not meaning
A well-formed filename can carry useful operational information: a project code, date, camera roll, scene and take, version number, cutdown or delivery specification. It can help a conform editor identify source material and help a producer distinguish a master from a social version.
What it rarely describes is everything happening inside the file.
A filename does not say that a contributor discusses sustainability halfway through an interview. It does not record that a cyclist crosses the background during the final seconds of a take, or that the client requested a warmer grade in the review attached to a particular version. It cannot contain every location, object, action, spoken phrase and possible use that another person may later recognise in the material.
Traditional file search therefore asks the user to translate a human memory into storage information. To find the shot, they must already know something about where it lives: the project, folder, date, camera, filename or metadata field. When none of those details are remembered, the search becomes manual viewing.
Meaning-based search reverses the direction. It starts with what the person actually knows: what was said, what appeared in frame, what the moment resembled or how it related to the project.
The limits of perfect organisation
Studios have good reasons to invest in naming conventions, metadata, logging and archive discipline. Those practices make media safer to move, easier to hand over and more reliable at conform and delivery. Meaning-based search does not replace them.
It addresses the part they cannot reasonably capture.
To describe a clip completely, a logger would need to anticipate every future question. A wide exterior might be relevant because of the building, the weather, the crowd, the camera movement, the wardrobe, a vehicle in the background or its resemblance to a new treatment. An interview answer might later matter for its exact wording, its broader subject, the person speaking or the project decision it influenced.
More tags can improve retrieval, but the possible descriptions grow faster than any practical tagging scheme. Logging also happens at a particular moment, for a particular production need. The description that helps build a first assembly may not be the one that helps a producer find material for a cutdown two years later.
The useful model is additive: preserve strong filenames and metadata, then make the content of supported media searchable as another layer.
Searching the spoken word
Speech is one of the clearest places to begin because editors often remember language more accurately than storage details.
An editor may recall the phrase “we wanted the space to feel open” without remembering who said it or on which interview day. A producer may need every interview in which sustainability is discussed, even when different contributors use words such as reuse, materials, waste or long-term impact. A post-production supervisor may want to locate the take where the client says “make it warmer” before checking how the note travelled into later versions.
When eligible media is transcribed, exact transcript search can locate the words that were actually captured. Semantic search can go a step further by looking for passages that express a related idea, even when the wording is not identical. The distinction matters. One is useful when the remembered quotation is reliable; the other is useful when the person remembers the subject rather than the sentence.
Neither should be treated as infallible. Accents, overlapping speakers, background noise, specialist vocabulary, multiple languages and poor recordings can all affect a transcript. Search results should therefore lead back to the media, with a timestamp and enough surrounding text for an editor to verify the moment quickly.
The goal is not to turn an imperfect transcript into unquestioned truth. It is to make spoken material navigable.
Searching what is in the frame
Some searches begin with no words from the footage at all.
Show me drone shots over water at golden hour. Find wide shots of an empty platform. Which versions still contain the old end card? These requests describe visible content rather than file properties.
Visual search can help identify objects, settings, actions, shot characteristics and broad features of a scene across locally generated previews. It may surface frames containing a shoreline, a person entering a doorway, a close-up of a particular product or an end card with the older design.
This is useful, but it is not the same as replacing an editor’s eye. A system may find scenes with similar visible ingredients while missing the performance, rhythm or narrative function that makes one shot right for the cut. It can narrow a large search space and reveal candidates. The creative judgement still belongs to the person reviewing them.
That limitation should shape the interface. Results are better presented as inspectable thumbnails with source details than as a single supposedly definitive answer.
Searching with an image
Sometimes the clearest query is another image.
A director may provide a still from a treatment. An editor may take a screenshot from an old assembly. A producer may have a moodboard image that captures the composition of the material they need. Rather than translating the reference into tags, they can use it to look for visually similar moments in the accessible media library.
There are several different tasks hidden inside that idea. Finding an exact or near duplicate is relatively concrete. Finding a similar composition — a solitary figure framed against a large architectural space, for example — is broader. Finding a similar type of scene is broader again. Searching for the same mood or aesthetic becomes more subjective, because those qualities depend on context, movement, performance, grade and the intention of the edit.
Image-led search is strongest when it remains honest about these differences. It can reveal footage that shares visible characteristics with a reference. It cannot be assumed to understand the complete artistic intention behind that reference.
A search, from question to source
Fictional data · generated exampleThis simulated conversation moves through filename, transcript, meaning and visual search, then uses a reference image to find related footage and trace one result back to its source.
Search modes are strongest together
Filename search, metadata search, transcript search, semantic search and visual search are not competing versions of the same feature. They answer different parts of a question.
A filename may establish that a clip belongs to the right camera roll. Metadata may confirm the shoot date and codec. A transcript may locate the sentence. Visual search may identify the matching setting. Semantic search may connect a loosely remembered idea to several differently worded interviews.
In practice, the most useful query often combines these signals. Find interviews from the launch project where sustainability is discussed, but only the selects used in the first assembly. The spoken subject identifies candidate moments; the project association narrows the material; the edit history changes which results matter.
This is the shift from finding files to finding evidence. The user should not need to decide in advance which technical search mode contains the answer. They should be able to describe the production need, then inspect how the results were assembled.
Project context changes the answer
Media does not acquire its full meaning from pixels and sound alone. A broad archive search can surface candidates first, then narrow them to a particular project before the editor inspects the source behind the most useful result.
The same shot may be a rejected option in one version, an approved hero moment in another and a rights concern in a third. A clip containing an old end card is only relevant if the studio knows which versions are still active. A shot remembered from the first treatment may have been replaced during production, restored during review or used only in a particular cutdown.
That is why useful footage search reaches beyond the media library. The answer may depend on a treatment, task, review comment, version history or production conversation. Search becomes more precise when these sources can provide context without being collapsed into one undifferentiated repository.
Foma is designed as that connective layer. It works across the studio’s existing media, documents, project tracker, review platform and team communication. It does not attempt to replace the MAM, DAM, archive system or NLE. Those systems remain responsible for the work they organise; Private Studio Intelligence helps the team ask questions across them.
This also makes ambiguity visible. If a review comment refers to “the warmer version” but two exports could match, the right response is not to invent certainty. It is to show the likely candidates and the evidence attached to each.
Results need provenance
A grid of plausible images is not enough for production work. Every useful result needs a route back to its source.
For footage, that may include the original file or storage path, project association, timestamp, transcript fragment and a locally generated preview. For a version question, it may also include the connected review comment, task or document that explains why the media matters.
Provenance lets an editor open the right moment, a producer verify the project relationship and a supervisor understand whether the result is current. It turns search from a suggestion mechanism into a practical path through the work.
It also supports correction. If the transcript is wrong, the thumbnail is unrepresentative or the system has connected the wrong project, the user can see the underlying evidence and make their own judgement. Private Studio Intelligence should shorten the route to verification, not ask people to trust an opaque answer.
Why local processing matters
Production archives can contain unreleased campaigns, confidential interviews, licensed assets and material subject to client or contractual controls. Making that content searchable should not quietly assume that full-resolution media will be sent to an external service.
A local-first approach changes the practical shape of the system. Supported media can receive thumbnails and previews generated on-site. Eligible media can be transcribed locally. Search can operate over those representations while the source material remains governed by the studio’s existing storage and access arrangements.
Deployment still needs to reflect the infrastructure and policies of the particular studio. Local-first is not a universal claim about every possible configuration. It is an architectural preference: keep media processing close to the archive, minimise unnecessary movement of sensitive material and preserve control over how sources are connected.
Read-only access is equally important. Searching for a moment should not rename a clip, move a folder, alter a task or change an approval. The intelligence layer helps people understand the existing production environment without taking authority over it.
From finding files to finding moments
Studios will continue to need disciplined filenames, metadata and storage structures. Those systems are how media remains operable over time.
But the questions people ask are richer than the fields available in a file record. They remember a sentence, a face, a composition, an action, a decision or the role a shot played in an earlier cut. Search should meet them at that level, then return the operational detail required to use the result safely.
The change is subtle but consequential. Instead of searching for the label attached to a file, the studio can search for the moment inside it — and the context around that moment.
That is what The whole studio, in one conversation means in practice: not a replacement for the production stack, but a shared way to ask what the studio already knows and follow the answer back to its sources.
Map your studio workflow
Start with three questions about footage that your team regularly asks other people because filenames and folders cannot answer them.
Tell us how your studio works