CBS 2026 Avatar Co-Host

“Hi Lisa, how are you?” — a named agent appears on a stage display, welcomes guests and speakers, and works with Mike live. What to call it, how to build it as a reusable app, and whether the tech is ready.

← Agentic Commerce  ·  CBS GTM  ·  OpenClaw  ·  Anam avatars  ·  Voice  ·  ElevenLabs  ·  Multi-agent harness  ·  Multiplayer agents  ·  Headless Empire OS  ·  Stage Agent product brief

Discussion open (this is not a commitment yet)

Cross Border Summit (annual conference, ~November) is a few months out — enough time to prototype and stage-rehearse, not enough to invent a flawless general co-host. The innovative move is real; the smart move is a bounded co-host with hard human control, not an unsupervised “AI MC for three days.”

Updated vision from the team: reuse (or extend) personas we already run in OpenClaw — especially Lisa Yuson ( lisayouson.com ) with established image + voice — and make the experience feel as simple as calling her by name on a dedicated front-of-stage display.

This page is the living research brief: product naming, reusable framework, OpenClaw vs stage web app, multi-persona roster, architecture, risks, and path to showtime.

The “Hi Lisa” moment (target experience)

  1. Dedicated screen stage-left / center (not buried in a slide deck only).
  2. Mike (or a clear wake phrase): “Hi Lisa, how are you?” / “Lisa, welcome our next speaker.”
  3. Lisa’s face appears (or animates), greets, runs a short segment (welcome, intro, recap).
  4. Optional short back-and-forth with Mike — then she yields the stage cleanly.
  5. Operator can always kill mute / freeze / cut to fallback.

That is not science fiction in 2026. Pieces exist today (STT, agent brain, TTS, talking avatar). What is still hard is reliability under stage conditions (latency, Wi‑Fi, wrong words, uncanny). So we design for called segments, not free-roaming co-host for eight hours.

What do we call this?

Useful to separate category name (industry language) from product name (what we ship) and persona name (who appears).

Layer Suggested language Example
Category Stage agent / live event co-host / embodied stage agent Blog research, RFPs, sponsor decks
Product / framework Stage Agent Runtime (working name) or Showhost / StagePresence Reusable web app + API we run at CBS, GFA meetups, Horizon booths
Persona Existing OpenClaw characters Lisa Yuson, Janice, … — not the product name
Event instance CBS 2026 configuration Vault + run-of-show + which personas are allowed on stage

Avoid calling the whole product “OpenClaw” (that’s the brain/harness) or only “avatar” (that’s just the face). The valuable thing is the stage loop: wake → brain → voice → face → human gate.

Recommended internal codename for now: stage-agent (framework) + persona lisa (and others) as plugins. Public brand can wait until a demo works twice.

Lisa Yuson and the OpenClaw roster — pull them in?

Yes — as the character and brain identity, not as “the whole stage product.”

  • Lisa Yuson ( lisayouson.com ) already has story, image, and voice investment — audiences and team may already “know” her from outreach / GTM ( CBS GTM already references Lisa in the agent stack).
  • OpenClaw is a strong candidate for the brain and tools (setup, sessions, skills) — same as Anam+OpenClaw on /anam.
  • Stage still needs a Presence app (fullscreen browser on the display): OpenClaw does not by itself put a 10-foot lip-synced face on a projector with a kill switch.

Single persona vs multiple

Approach Pros Cons
One face for CBS (Lisa) Clear brand, simpler rehearsal, one voice model to stress-test Less variety; if Lisa fails, whole bit fails
Roster (Lisa welcomes, another does recaps…) Role-first (matches Hermes / multiplayer rules) More assets, more failure modes, more operator complexity

v1 recommendation: one primary stage persona (Lisa or a dedicated CBS character). Keep other OpenClaw agents for off-stage ops (GTM, ops room). Multi-persona on stage only after one persona is boringly reliable.

Is this a web app / reusable framework?

Yes — that should be the product shape. Not a one-off Zoom link. A small Stage Agent Runtime we can re-use at: CBS, GFA meetups, Horizon AI booths, webinars, sponsor dinners.

What already exists (buy / assemble)

  • Talking avatars: Anam, HeyGen Interactive Avatar, similar real-time face APIs
  • Voice: ElevenLabs, Cartesia, etc. (notes)
  • Agent brains: OpenClaw, Hermes, Grok Build, Claude/Codex, QM-style company harnesses
  • Live STT: Whisper-class streaming, Deepgram, vendor SDKs

Nobody sells “Mike says Hi Lisa and the PA + projector just work at a conference” as one polished product. The gap is the orchestration shell for events — we either assemble it or build a thin framework.

What we would own (the framework)

stage-agent/   (working product name)
├── presence/          # Fullscreen web app for the stage display (avatar canvas)
├── console/           # Operator tablet: next line, green light, kill, segment picker
├── gateway/           # Wake word / push-to-talk / host mic routing
├── brain-adapter/     # OpenClaw | Hermes | HTTP agent — same interface
├── persona-packs/     # lisa/, cbs-host/, …  (voice id, face id, soul.md, vault)
├── event-vault/       # agenda, bios, sponsors, rulings (HEOS-shaped)
└── fallbacks/         # Pre-rendered mp3/mp4 per segment if network dies

The presence + console + kill switch + vault are the reusable core. Persona packs and event vaults swap per show. Brain can stay OpenClaw for Lisa continuity.

Minimal viable web experience

  1. stage.example.com/display/lisa — kiosk/fullscreen on the stage monitor (Chrome, autoplay unlocked via operator click once).
  2. stage.example.com/console — operator only (auth): “Welcome guests” / “Intro speaker X” / “Summarize last 3 min” / Kill.
  3. Mike’s mic → STT → optional wake “Lisa” → brain (OpenClaw session with CBS vault) → TTS + avatar stream on display.
  4. Confidence line on console: what she will say before (or as) she speaks, depending on latency budget.

Is the technology too early?

Capability 2026 readiness for stage
Branded character + voice we already love Ready (Lisa pack)
Scripted welcome / intro / sponsor read Ready if vault-locked
Called-on interaction with host (“How are you?” + short bit) Ready with rehearsal — expect 1–3s lag; write bits that tolerate it
Long free-form co-hosting, comedy, crisis MC Early / risky — don’t bet the show
Perfect venue networking Always the weak link — plan offline fallbacks

Bottom line: working toward “Hi Lisa” is exactly the right ambition for CBS. Shipping a reliable segment co-host is on the table; shipping a full AI co-MC is not required to innovate or wow the room.

Is it too much to ask?

Scope Verdict
Full autonomous MC for multi-day agenda, unscripted Q&A, jokes, crisis control Too much for a first year — latency, hallucination, brand risk
Called-on co-host: Mike triggers segments; avatar speaks prepared or tightly grounded replies; human always can cut Realistic innovator — same energy as CBS innovation, controlled blast radius
Segment specialist: only intros, only session summaries, only sponsor reads, only “what just happened” recaps Best first ship — demo wow + recoverable failures

Recommendation: ship as a named persona (strong default: Lisa Yuson, already character-complete) inside a reusable stage-agent web runtime, with a visible kill switch, a stage operator, and a shortlist of moments — not free roam.

What “agent co-host” means on stage

Three products must work together in under ~2–4 seconds perceived latency when possible:

1. Face (avatar)

Consistent character image: lip-sync, optional gestures, big-screen safe. Options: Anam.ai, HeyGen live avatars, custom Unreal/Unity, or simpler “voice + fixed portrait / looping video” for v1.

2. Voice

Fixed branded voice (clone or designed). Low latency TTS / streaming. See ElevenLabs, VibeVoice, voice agent notes.

3. Brain

Agent with CBS knowledge vault: agenda, speaker bios, sponsor copy, house rules, “never say” list. Harness: OpenClaw / Hermes / Grok Build / company stack — not raw ChatGPT on Wi‑Fi alone.

Mike (or stage manager) is the fourth layer: human gate — push-to-talk, segment select, interrupt, or “read this script only.”

Jobs the co-host can own (CBS-shaped)

Moment Agent role Risk
Open / close day Welcome, house rules, “who’s in the room,” safety/exit notes Low if scripted
Speaker intros Bio from vault, 30–45s, correct titles, no improvised resume Med if bio wrong
Session summarize 3 bullets after a talk (notes + transcript feed) Med — need good capture
Sponsor / partner reads Exact approved copy only Low if locked text
Audience Q filter Cluster questions; surface top 3 for Mike Low (assist only)
Logistics (“lunch is…”, “track B is…”) From live ops board / HEOS-style TODAY for the event Low if ops feed is clean
Banter with Mike Short call-and-response, rehearsed bits High if free-form
Live debate / unscripted expert answers Defer to humans Too high for v1

Architecture (discussion draft)

┌──────────────────────────────────────────────────────┐
│ STAGE                                                 │
│  Mike (mic)  ←→  Operator tablet  ←→  Kill switch     │
│       ↓                                               │
│  Confidence monitor (what agent will say next)        │
│       ↓                                               │
│  Big screen: Avatar video stream                      │
│  PA: Voice TTS                                        │
└──────────────────────────────────────────────────────┘
              │ Wi‑Fi / hardline preferred
              ▼
┌──────────────────────────────────────────────────────┐
│ CONTROL PLANE (backstage laptop / mini)               │
│  • Segment buttons: Intro | Summarize | Sponsor | Ops │
│  • Script lock vs “open question” mode                │
│  • Latency / fallback audio file library              │
└──────────────────────────────────────────────────────┘
              │
              ▼
┌──────────────────────────────────────────────────────┐
│ BRAIN                                                 │
│  CBS vault: agenda, bios, sponsors, rulings           │
│  Agent harness (OpenClaw / Hermes / Grok / QM-style)  │
│  Optional: multiplayer room for ops team              │
│           (GFAVIP Webchat / Slack)                    │
└──────────────────────────────────────────────────────┘
              │
              ▼
┌──────────────────────────────────────────────────────┐
│ RENDER                                                │
│  Avatar provider (Anam / HeyGen / simpler v1)         │
│  TTS provider (ElevenLabs / etc.)                     │
│  Optional STT of Mike for “call by name”              │
└──────────────────────────────────────────────────────┘

Same layers as our multi-agent harness map: surface (stage), harness (brain), vault (CBS knowledge) — plus a broadcast renderer.

Human control model (non-negotiable)

  • Mike calls the co-host — not continuous free talk unless rehearsed.
  • Operator confirms next utterance on a confidence screen (or green-lights segment).
  • Hard kill — mute TTS + freeze/hide avatar in <1s (hardware switch if possible).
  • Fallback packs — pre-rendered audio/video for intros and sponsors if network dies.
  • Rulings file — words never said, legal, sponsor constraints (same idea as knowledge base rulings).
  • Permissions — no live send of emails/posts from stage agent without offline review.

Stack options (research shortlist)

Layer Options we already document / use Notes for stage
Avatar Anam + OpenClaw, HeyGen, Hyperframes-style video-as-code later Prefer providers with streaming low-latency; stress-test on venue Wi‑Fi
Voice ElevenLabs, voice agents, VibeVoice Clone or design a CBS character, not Mike’s exact voice (confusion / uncanny)
Brain OpenClaw, Hermes, Grok Build, HEOS vault Offline-capable brief + online model; never only cloud chat in browser
Ops surface GFAVIP Webchat, multiplayer room, stage tablet Ops posts agenda changes; agent reads same source of truth
Capture for summaries Room mics → STT → agent summarize Consent + recording policy for speakers/attendees

v0 (lowest risk): fixed character art + excellent TTS + teleprompter-style agent-generated script on operator screen that Mike or co-host reads — still “AI assisted,” less avatar risk. v1: live avatar + locked segments. v2: open Q with heavy grounding after v1 works twice in dress rehearsal.

CBS knowledge vault (what the agent must know)

Treat as a mini HEOS / knowledge base for the event:

  • agenda.md — day, room, times (single source; ops owns updates)
  • speakers/*.md — approved bios only
  • sponsors/*.md — approved reads, logos, never-claims
  • voice.md — co-host tone (warm, cross-border, not salesy, not political)
  • rulings.md — never improvise legal, medical, investment advice; defer to Mike
  • run-of-show.md — when avatar is allowed on, cues, music beds

Fill with reverse-prompting style interviews of the production team (Alex Finn reverse prompting) so the vault is complete before tech build.

Risks & mitigations

Risk Mitigation
Venue Wi‑Fi dies Hardline or local hotspot; offline fallback media; LTE backup
Latency kills banter Segment-based not free conversation; pre-generate where possible
Wrong bio / sponsor claim Script lock from vault; operator preview; no free web search on stage
Uncanny / creepy brand hit Friendly character design; short segments; humor that doesn’t mock speakers
Recording / likeness rights Don’t clone a real person without contracts; original avatar preferred
Demo becomes the story (failure) Position as experiment; Mike always can finish alone

Path to November (rough timeline)

  1. Now (discussion): this page — agree scope (v0/v1), persona brief, kill-switch owner.
  2. +2–3 weeks: vault draft (agenda shape, sample bios, voice/rulings); pick avatar+voice vendor for a lab demo.
  3. +6 weeks: closed-room demo: Mike calls 3 segments (intro, summary, sponsor) with operator gate.
  4. +10 weeks: dress rehearsal in similar room (PA, projection, latency); fallback pack ready.
  5. Show week: limited live moments only; record for GTM/content after.

Parallel: CBS GTM and attendee systems stay human-led — GTM plan, GFAVIP perks — co-host is production innovation, not a substitute for ops.

Open decisions (for the team)

  1. Primary persona: Lisa Yuson vs dedicated CBS-only character vs multi-roster later?
  2. Product codename for the reusable app (stage-agent / Showhost / other)?
  3. Brain: OpenClaw (Lisa continuity) vs separate low-latency stage brain with same soul files?
  4. v0 (voice + still/looping portrait) vs v1 (live lip-sync avatar) for first public use?
  5. Who is stage operator / kill-switch owner?
  6. Which moments are guaranteed live vs pre-rendered backup?
  7. Recording + release rights for co-host segments (and Lisa brand guidelines)?
  8. Budget: avatar SaaS + voice + on-site hardware + engineer days?

Related on this site

Anam + OpenClaw

Real-time talking avatars + agent brain — face layer for stage presence.

Read →

OpenClaw

Likely brain for Lisa continuity — sessions, skills, fleet.

Read →

CBS GTM 2026

Human + agent GTM for the summit — separate from stage co-host but same event.

Read →

Multi-agent harness map

Where stage brain sits among QM, Hermes, OpenClaw, GFAVIP rooms.

Open map →

Multiplayer agents

Ops team + co-host agent sharing context in a room.

Read →

Knowledge base / HEOS

Vault + rulings pattern for agenda, bios, sponsor copy.

KB →   HEOS →

Horizon AI / booth

Other AI presence at events in our notes.

Horizon →

Voice stack

ElevenLabs, voice agents, TTS choices.

Voice →   ElevenLabs →

Reverse prompting

Interview production team to fill the CBS vault.

Read →

Bottom line

Yes — work toward “Hi Lisa” on a dedicated stage display. The tech for a called co-host is ready enough if we treat this as a reusable stage-agent web runtime (presence + console + vault + kill switch), plug in OpenClaw personas like Lisa as packs, and bound the show to rehearsed segments. Not too early to innovate; too early to let the agent run the conference unsupervised.

Product brief: Stage Agent Runtime · Agent build handoff: procedures/stage-agent-build-brief.md (pass into Grok Build /goal — Phase 0 scaffold first).

Comments

Approved comments appear below. Log in once with GFAVIP — it applies across the whole site. GFAVIP login

View comments archive