Skip to main content

Production and QA

How We Capture, QA and Version Human Video Datasets for AI

A practical account of how we turn an AI dataset brief into a traceable package through written protocols, take-level quality assurance, manifests, checksums, package freezes and controlled delivery.

Author
Continental People Production Team
Reviewed by
Georgii Ilin, Founder & CEO
Published

Human video data is easier to evaluate and integrate when each delivered file can be traced to an agreed capture requirement, an acceptance decision and a package version. Our workflow starts with a written protocol and ends with a frozen, privately delivered package—not simply a folder of clips.

The exact checks and delivery artifacts depend on the dataset. A synchronized multi-view action package, a repeated facial-motion protocol and an archival collection do not have identical evidence structures. We describe those differences instead of presenting one generic QA claim for every product.

The workflow at a glance

Brief → protocol → capture → take-level QA → targeted reshoots → package reconciliation → version freeze → recipient review and private delivery

For an off-the-shelf dataset, the brief is an internal product specification. For a custom capture, it begins with the buyer's intended model workflow, required actions, participant profile, recording conditions, metadata needs and acceptance criteria.

Start with intended use and acceptance criteria

Before recording, we translate the brief into a capture protocol. Depending on the project, it can define:

  • prompts or actions and their required start and end states;
  • variants, repetitions and relevant failure cases;
  • environments, objects, wardrobe and continuity requirements;
  • camera views, resolution, frame rate, synchronization and audio;
  • metadata fields and the identifiers that connect files to the protocol;
  • participant-permission requirements that affect capture or delivery;
  • technical, content and coverage acceptance rules.

For novel or technically demanding work, a pilot can be agreed before full production. The pilot is used to confirm framing, performance, metadata and acceptance logic while changes are still inexpensive. If an approved scope changes, we document the change rather than silently folding it into the dataset.

Keep events, views and files connected

A file count is not always an event count. In multi-view work, two cameras may record the same physical action. We map the files to the same event and distinguish the view, so the same action is not mistaken for two independent events.

Repeated-prompt captures need a different mapping. Metadata connects each clip to the relevant participant, capture condition, prompt and repetition. Where participant identifiers are required, client-facing package records use pseudonymous production identifiers. Legal names, signed participant releases and other direct identity records remain outside the client media tree and are managed separately.

Media and metadata can also have different revision needs. In our FaceMotion package structure, source media is treated as an immutable media revision while attribution and QA metadata can be revised separately. This allows a documented metadata correction without silently replacing the original recording. If media itself must be replaced, the affected package and integrity records require a new revision.

Apply technical, content and coverage QA

Acceptance criteria vary by project, but we separate at least three questions:

QA layerQuestionTypical checks
Technical QAIs the file technically usable under the agreed specification?File readability, container and codec, resolution, frame rate, focus, exposure, framing, synchronization and audio where required
Content QADoes the recording contain the requested material?Correct action or prompt, start and end state, variant, repetition, view, visibility and performance
Coverage QAIs every required protocol slot accounted for?Expected combinations, accepted takes, rejected takes, missing items, deviations and reshoot status

A technically valid file can still fail content QA. A correct performance can still fail because a required body part is out of frame or synchronized views cannot be matched. Rejected or ambiguous takes do not count toward required coverage. Where the production context can still be reproduced, rejected or ambiguous protocol slots are flagged for targeted reshoot.

Known limitations remain part of the delivery record. Hiding a deviation may make a package look cleaner, but it prevents a buyer from evaluating the data accurately.

Reconcile the package against its structure

Different package designs require different reconciliation rules.

In Everyday Dexterity & Packaging (EDP) v1, 40 physical events reconcile to 80 original files because every event has two synchronized camera views. The event count is therefore 40, not 80. For the same reason, adding the duration of both synchronized views does not produce the unique action timeline.

In the FaceMotion capture protocol, 32 prompts × 2 lighting conditions × 2 repetitions define 128 expected clips for a complete participant capture. Coverage QA checks the expected prompt-condition-repetition slots, not just the raw clip count.

Archival material requires a different statement. We do not describe an archival collection as if it had been captured under our current protocol. Its evidence is scoped to what can actually be verified: media inventory, duration, audio state, available metadata, documented limitations, provenance and applicable rights records.

Freeze a version that can be reconciled and verified

Working files are not a delivery version. Before controlled delivery, an approved media inventory is separated from ongoing production work and identified as a specific package version.

Depending on the product and order, the package record can connect:

  • the approved media inventory;
  • a manifest or metadata table;
  • a package and metadata version;
  • a README or data dictionary;
  • known limitations or a change record;
  • a checksum inventory and QA summary where specified;
  • the applicable order, license and delivery record.

A manifest describes what should be present and how files relate. A checksum is a file fingerprint that can confirm that a file has not changed since that checksum was generated. Neither one proves that a label is correct, that the applicable permissions are sufficient or that the dataset will improve a particular model.

If files are added, removed or replaced after a freeze, the result is not silently presented as the same package. The version record, manifest and any affected integrity artifacts must show what changed.

Verify without publishing private storage

Package verification does not require public bucket paths or permanent download URLs. Before delivery, we apply the verification steps defined for that package. Depending on the applicable product gate, these can include reconciling file counts, structure, metadata and checksums against the frozen package record.

Private storage locations, object keys and participant records are not public product documentation. Public pages describe the package and its availability; they are not delivery endpoints.

Deliver under a recorded order and license

Our current delivery process is manual. We review each request, record the applicable order and license acceptance, and then provide approved access privately. The applicable product version and order identify the delivery contents for that SKU or project.

Participant releases and direct identity records are not placed in the delivered media tree. The buyer's rights and restrictions come from the applicable dataset license and order documentation, not from possession of private participant paperwork.

What this evidence does—and does not—show

Our QA and version records support package traceability, coverage review and file-integrity checks. They help a buyer understand what was requested, what was accepted and which version was delivered.

They are not an independent audit, a legal opinion or a guarantee of downstream model performance. QA records whether the package passed the checks defined by the applicable capture and package specification. A buyer should still evaluate the sample, specifications, rights and known limitations against its own intended workflow.

Frequently asked questions

What is a human video dataset manifest?

A manifest is a structured inventory of the package. It can identify expected media files and connect them to events, views, prompts, participants, conditions or other metadata defined by the dataset.

Why can the file count be higher than the event count?

One physical event can produce several files. A synchronized two-view capture, for example, produces two camera files for one event. Event count and file count should therefore be reported separately.

What does a checksum prove?

A matching checksum confirms that the file being checked is byte-for-byte identical to the file for which that checksum was generated. It does not prove content quality, label accuracy, legal clearance or suitability for a model.

Are signed participant releases included in the media package?

No. Signed releases, legal names and other private participant records are kept outside the client media tree. The applicable buyer license and order define the permitted use of the delivered package. If a buyer requires deeper review of release coverage, we address it separately and case by case without placing private participant documents in the media tree or publishing them.

Planning a new capture? Use the custom human video dataset brief to organize requirements before scoping.

For the rights-review side of a package, use the model releases buyer due-diligence checklist.

From requirements to a controlled package

Have acceptance criteria already? Send us the intended use, actions, participant profile, views, metadata and delivery constraints. We can turn them into a capture and QA plan.