Created Last updated

Decision post · James Meadlock & Milo · native Hermes capability

Video understanding without a pipeline

Measured update, August 4: Tonbi's 420-segment captions passed; short direct-video canaries passed; the full 17-minute inline route failed semantic verification. Google's recommended Files API route then recovered a usable full-video answer, with one page-count error corrected during QA. Roxy remains separately gated.
Caption ingestPASS
Short videoPASS
Long inline videoFAIL
Long Files APIPASS*

The first plan treated video like a new corpus product: acquire media, transcribe it, build timestamped nodes, index it, deduplicate it, and eventually give it a database. That architecture is defensible if the job is to operate a video library. It is unnecessary if the job is usually, "understand this video and answer my question."

Hermes already has most of the useful primitives. The new design starts there and refuses to create infrastructure until real use proves something is missing.

The routing rule

Three-tool native Hermes video-understanding architecture A question and source are classified into YouTube captions, canonical web text, or a short direct video path now validated through native Gemini in Milo. Evidence is merged into a cited answer, with optional human-readable retention. URL or file + a real question classify source • identify evidence needed youtube-content public captions + timestamps v1 • dependency canary web_extract official pages • transcripts • PDFs v1 • text before pixels video_analyze short permitted direct clips SHORT GREEN • LONG FAIL Evidence-aware answer real timestamp • page citation • model-observed label

Short controlled direct-video canaries are green in Milo. The 17-minute inline Tonbi canary failed semantic verification even though transport succeeded. A separate provider-native Files API probe passed with a QA correction, but it is not yet part of Hermes video_analyze; Roxy remains separately gated.

SourceFirst toolWhat it can support
YouTube with public captionsyoutube-contentSpoken claims, chapters, quotes, and timestamp links
Official event page, transcript, deck, or PDFweb_extractCanonical page or document evidence; no invented video timestamps
Direct permitted clip under the measured limitvideo_analyzeModel-observed visual action and broad audio events; never a verbatim transcript
Gated, DRM, captionless, oversized, or dynamic-player sourceStopReport the limitation and request a permitted artifact

What is proven today

PrimitiveMeasured stateCall
web_extractPassed canonical HTML and official PDF canaries against Apple and Manulife sourcesGreen in Milo
youtube-contentDependency installed in the Hermes runtime; provenance regression suite passed 3/3 and a live timestamped transcript canary passedGreen in Milo
video_analyzeShort controlled URL and local-file canaries passed. The full 17-minute inline Tonbi run was accepted but hallucinated unrelated content. A separate provider-native Files API probe recovered a useful answer with QA corrections.Inline long fails; Files API passes*

A real YouTube ingestion: Tonbi

We ran the text-first route against Tonbi's Feeding GBrain Your Emails Autonomously with Hermes Agent.[6] The result was a complete 17:06 English transcript with 420 timestamped segments. YouTube reported the captions as auto-generated, so the timestamps are useful evidence while the wording is not treated as verbatim.

FieldIngested result
Routeyoutube-content; no downloader, media capture, or direct-video model
ProvenanceEnglish, auto-generated captions; 420 segments; 17:06
Core patternDeterministic script collects; agent applies judgment; the knowledge system remembers
Security boundaryAllowlisted senders expose content; unknown senders remain in human review; existing entities are updated but new ones are not created autonomously
ArtifactFull timestamped transcript retained locally for audit and re-use

The ingestion recovered the workflow cleanly:

Evidence limit: this proves caption ingestion, provenance handling, and grounded timestamp links. It does not prove that Gemini inspected Tonbi's slides or other visual details; a normal YouTube watch page is not passed to video_analyze as if it were a direct media file.

Native-video follow-up: accepted transport, failed semantics

At James's request, we then made an explicit one-off exception to the provisional direct-video envelope and ran the full locally acquired media rendition through Hermes video_analyze. The file was 20.47 MiB, 640×360, and 17:04.163. That is shorter than YouTube's displayed 17:06 and outside both initial guardrails: under three minutes and under 15 MiB. It was not represented as a short excerpt.

Measured fieldFull Tonbi result
Hermes transportsuccess:true; first completion returned in 18.09 seconds
Semantic acceptanceFAIL — the first answer recast Tonbi as a CI/security product and invented GitHub Actions, PR gates, tonbi.yaml, and a CLI scanner
Constrained retryFAIL — it recovered the broad topic only after being told it, then invented second-long scene timings, collector.py, JSON fields, commands, and terminal logs not present on screen
Independent frame checkConfirmed the actual slides and terminal result below; these observations are QA evidence, not claims produced reliably by the full-video model run

Independent frames showed what the model should have reported:

Inline acceptance verdict: the full inline Tonbi run is a useful negative canary. Gemini consumed enough video to incur a large media-token bill, but neither success:true response was safe to publish as visual evidence. Captions remain the production route for this source; bounded clips remain the accepted Hermes-native lane, while the Files API result below is a reviewed provider-native escalation rather than a wired Hermes capability.

Measured token usage and cost

Hermes's current video_analyze result exposes the completion but not provider usage metadata. To measure rather than guess, we sent the same full file and a constrained prompt through the same native Gemini 3.6 Flash endpoint. Google reported 93,187 video tokens plus 141 text-input tokens, 95 candidate tokens, and 2,400 thinking tokens. That metered call took 21.415 seconds. Its output was truncated and was not used as content evidence.

Billing componentMeasured tokensStandard rateCost
Input93,328$1.50 / 1M$0.139992
Output, including thinking2,495$7.50 / 1M$0.018712
Measured native call95,823 totalStandard tier$0.158704, or about $0.159

Those rates come from Google's official Gemini 3.6 Flash standard-tier pricing.[7] Because the Hermes wrapper did not expose usage for its two attempts, their exact charges are unavailable. The same-file measurement puts each attempt at roughly $0.16–$0.18, depending on generated and thinking tokens. Including the two Hermes attempts and the one metered native call, this investigation cost approximately $0.48–$0.52. That total is explicitly an estimate; $0.158704 is the measured reference call.

Files API follow-up: usable after QA

The first full-video experiment used inline bytes because that is what the patched Hermes adapter emits. Google's current guidance points elsewhere for this source: inline data is intended for short, one-off video, while the Files API is recommended when duration is significant or the same media will support multiple prompts.[5][8] We therefore uploaded the same 20.47 MiB, 17:04.163 MP4 through Gemini's resumable Files API, waited for its state to become ACTIVE, referenced its file URI in generateContent, and deleted it immediately after the response. A list-after-delete check found zero retained files.

Files API stageMeasured result
Upload3.435 seconds
Provider processing20.746 seconds
Gemini generation14.426 seconds
End to end38.640 seconds

The result was materially better than the inline runs. Without being given the answers, Gemini identified the real script collects → agent judges → brain remembers architecture; named the major visible slide sequence; described the collector's no-LLM/no-direct-write boundary; distinguished the structured page contract from the blob; recovered the demonstrated 1 allowlisted / 1 review / 4 noise dropped counts; and reported the three visible security gates. Those claims align with the caption track and sampled frames.

PASS*, not blind trust: the response also muddled the page-count transition as “29 to 32 / 30”; the visible terminal ground truth is 29 to 30 pages. It added small-terminal details that were not independently accepted. The route passed as a useful reviewed analysis, not as an autonomous citation generator.
Files API billing componentMeasured tokensStandard rateCost
Input93,331: 93,187 video + 144 text$1.50 / 1M$0.139997
Output1,683$7.50 / 1M$0.012623
Files API inference95,014 totalStandard tier$0.152619, or about $0.153

The Files API itself is available at no cost; inference tokens are billed normally.[7][8] This run reported no separate thinking-token count. The better route was slightly cheaper than the $0.159 metered inline reference because it produced 1,683 answer tokens without the inline probe's 2,400 thinking tokens. Added to the earlier inline investigation, cumulative measured/estimated spend for all four full-video calls was approximately $0.632–$0.672.

The video result required more than selecting a model name. Hermes emits an OpenAI-shaped video_url data URL. In v0.20.0, the native Gemini adapter translated image_url into Gemini inlineData but silently dropped video_url. We added a regression test, watched it fail, made the adapter treat both media types consistently, and then passed 9 native-adapter tests and 41 combined adapter/vision tests.[4][5]

Video pathLatencyObserved answer
Direct native Gemini API5.520 sWhite rabbit, butterfly, falling apple, no spoken dialogue; Google reported 910 video prompt tokens
Hermes video_analyze, direct URL7.248 sWhite rabbit, butterflies, meadow, apple, music, no speech
Hermes video_analyze, local file5.493 sWhite rabbit, butterfly, meadow, music, no speech

The negative matrix still matters. Copilot Gemini, Anthropic, xAI, OpenAI Codex, Nous Portal, and the tested local Qwen endpoint did not preserve or accept this payload. Several returned ordinary refusal prose while Hermes reported success:true. That boolean proves a completion arrived, not that the model saw the clip.

James/Milo and Bob/Roxy

James/MiloBob/Roxy
Default behaviorAnswer on demand; save a curated Obsidian note when usefulAnswer on demand inside the locked cited-research scope
YouTubeLive timestamped canary passed with reported caption provenanceCanary from Bob's actual hosted egress; no proxy workaround if YouTube blocks it
Web sourcesCanonical HTML and official PDF canaries passedSame path through the hosted Tool Gateway; hosted canary pending
Direct videoBounded video_analyze clips pass; the full inline run failed, while a separate provider-native Files API probe passed with QA correctionsOff by default; requires the adapter fix plus its own credential, privacy, cost, and semantic canary
StorageObsidian only when the answer deserves to surviveOrdinary document path after Bob's corpus exists; no media database
Still excluded: LlamaIndex, a Whisper or FFmpeg service, OCR workers, browser-driven media acquisition, task queues, a Forge video database, raw media in Postgres, and any attempt to bypass registration, CAPTCHA, DRM, or rights controls.

Evidence rules

The YouTube helper correction is complete: requested fallback languages stay separate from observed caption metadata, embedded snippet newlines are normalized, and three regression tests cover provenance and serialization behavior.[1]

When storage becomes justified

A database is not triggered by an arbitrary use count. The signal is repeated retrieval: the same source is queried again, one question spans several sources, or the same transcript would otherwise be fetched for a third time. A small permitted transcript cache may be enough. A structured corpus remains a separate decision.

Independent review

We sent the replacement plan and early failures to Claude Opus 5 through Nous. It returned APPROVE-WITH-CHANGES. The useful constraints survived: text first, honest caption provenance, a bounded direct-video envelope, Bob-hosted egress testing, separate profile activation, and retrieval-driven storage. The short direct-video lane is real in Milo; the Tonbi comparison shows that long inline video remains unsafe while the provider-native Files API is promising only with explicit QA. The separate Roxy gate remains.

Two reviewer claims did not survive verification. The failed automatic route was not running Opus 5, and the suggested local Qwen3-VL endpoint rejected Hermes's video_url payload with HTTP 422. Independent review is evidence, not authority.

What remains

  1. Upstream the tested video_url to Gemini inlineData adapter fix so a Hermes update cannot overwrite the local correction.
  2. Repeat the text-first canaries in Bob's actual hosted profile and egress.
  3. Keep long-form inline video outside the accepted envelope. Evaluate a narrow Hermes Files API integration with repeat canaries and mandatory semantic QA before treating it as a supported route.
  4. Keep Roxy's direct-video route off until the hosted build contains the fix and her own credential, privacy, cost, and semantic canary passes.
  5. Write one versioned shared operating skill and install reviewed, profile-scoped copies.
Definition of done: I can hand either agent a supported source, get an answer whose evidence granularity is honest, and watch it stop cleanly when the source falls outside the contract. No new service has to stay alive afterward.

Sources

  1. Hermes bundled youtube-content skill
  2. Hermes built-in tools reference
  3. Hermes model and auxiliary-model configuration
  4. Hermes video_analyze implementation
  5. Google Gemini video-understanding documentation
  6. Tonbi's AI Garage — Feeding GBrain Your Emails Autonomously with Hermes Agent
  7. Google Gemini 3.6 Flash pricing
  8. Google Gemini Files API documentation