CBS 2026 Avatar Co-Host
“Hi Lisa, how are you?” — a named agent appears on a stage display, welcomes guests and speakers, and works with Mike live. What to call it, how to build it as a reusable app, and whether the tech is ready.
← Agentic Commerce · CBS GTM · OpenClaw · Anam avatars · Voice · ElevenLabs · Multi-agent harness · Multiplayer agents · Headless Empire OS · Stage Agent product brief
Discussion open (this is not a commitment yet)
Cross Border Summit (annual conference, ~November) is a few months out — enough time to prototype and stage-rehearse, not enough to invent a flawless general co-host. The innovative move is real; the smart move is a bounded co-host with hard human control, not an unsupervised “AI MC for three days.”
Updated vision from the team: reuse (or extend) personas we already run in OpenClaw — especially Lisa Yuson ( lisayouson.com ) with established image + voice — and make the experience feel as simple as calling her by name on a dedicated front-of-stage display.
This page is the living research brief: product naming, reusable framework, OpenClaw vs stage web app, multi-persona roster, architecture, risks, and path to showtime.
The “Hi Lisa” moment (target experience)
- Dedicated screen stage-left / center (not buried in a slide deck only).
- Mike (or a clear wake phrase): “Hi Lisa, how are you?” / “Lisa, welcome our next speaker.”
- Lisa’s face appears (or animates), greets, runs a short segment (welcome, intro, recap).
- Optional short back-and-forth with Mike — then she yields the stage cleanly.
- Operator can always kill mute / freeze / cut to fallback.
That is not science fiction in 2026. Pieces exist today (STT, agent brain, TTS, talking avatar). What is still hard is reliability under stage conditions (latency, Wi‑Fi, wrong words, uncanny). So we design for called segments, not free-roaming co-host for eight hours.
What do we call this?
Useful to separate category name (industry language) from product name (what we ship) and persona name (who appears).
| Layer | Suggested language | Example |
|---|---|---|
| Category | Stage agent / live event co-host / embodied stage agent | Blog research, RFPs, sponsor decks |
| Product / framework | Stage Agent Runtime (working name) or Showhost / StagePresence | Reusable web app + API we run at CBS, GFA meetups, Horizon booths |
| Persona | Existing OpenClaw characters | Lisa Yuson, Janice, … — not the product name |
| Event instance | CBS 2026 configuration | Vault + run-of-show + which personas are allowed on stage |
Avoid calling the whole product “OpenClaw” (that’s the brain/harness) or only “avatar” (that’s just the face). The valuable thing is the stage loop: wake → brain → voice → face → human gate.
Recommended internal codename for now: stage-agent (framework) + persona lisa (and others) as plugins.
Public brand can wait until a demo works twice.
Lisa Yuson and the OpenClaw roster — pull them in?
Yes — as the character and brain identity, not as “the whole stage product.”
- Lisa Yuson ( lisayouson.com ) already has story, image, and voice investment — audiences and team may already “know” her from outreach / GTM ( CBS GTM already references Lisa in the agent stack).
- OpenClaw is a strong candidate for the brain and tools (setup, sessions, skills) — same as Anam+OpenClaw on /anam.
- Stage still needs a Presence app (fullscreen browser on the display): OpenClaw does not by itself put a 10-foot lip-synced face on a projector with a kill switch.
Single persona vs multiple
| Approach | Pros | Cons |
|---|---|---|
| One face for CBS (Lisa) | Clear brand, simpler rehearsal, one voice model to stress-test | Less variety; if Lisa fails, whole bit fails |
| Roster (Lisa welcomes, another does recaps…) | Role-first (matches Hermes / multiplayer rules) | More assets, more failure modes, more operator complexity |
v1 recommendation: one primary stage persona (Lisa or a dedicated CBS character). Keep other OpenClaw agents for off-stage ops (GTM, ops room). Multi-persona on stage only after one persona is boringly reliable.
Is this a web app / reusable framework?
Yes — that should be the product shape. Not a one-off Zoom link. A small Stage Agent Runtime we can re-use at: CBS, GFA meetups, Horizon AI booths, webinars, sponsor dinners.
What already exists (buy / assemble)
- Talking avatars: Anam, HeyGen Interactive Avatar, similar real-time face APIs
- Voice: ElevenLabs, Cartesia, etc. (notes)
- Agent brains: OpenClaw, Hermes, Grok Build, Claude/Codex, QM-style company harnesses
- Live STT: Whisper-class streaming, Deepgram, vendor SDKs
Nobody sells “Mike says Hi Lisa and the PA + projector just work at a conference” as one polished product. The gap is the orchestration shell for events — we either assemble it or build a thin framework.
What we would own (the framework)
stage-agent/ (working product name)
├── presence/ # Fullscreen web app for the stage display (avatar canvas)
├── console/ # Operator tablet: next line, green light, kill, segment picker
├── gateway/ # Wake word / push-to-talk / host mic routing
├── brain-adapter/ # OpenClaw | Hermes | HTTP agent — same interface
├── persona-packs/ # lisa/, cbs-host/, … (voice id, face id, soul.md, vault)
├── event-vault/ # agenda, bios, sponsors, rulings (HEOS-shaped)
└── fallbacks/ # Pre-rendered mp3/mp4 per segment if network dies The presence + console + kill switch + vault are the reusable core. Persona packs and event vaults swap per show. Brain can stay OpenClaw for Lisa continuity.
Minimal viable web experience
stage.example.com/display/lisa— kiosk/fullscreen on the stage monitor (Chrome, autoplay unlocked via operator click once).stage.example.com/console— operator only (auth): “Welcome guests” / “Intro speaker X” / “Summarize last 3 min” / Kill.- Mike’s mic → STT → optional wake “Lisa” → brain (OpenClaw session with CBS vault) → TTS + avatar stream on display.
- Confidence line on console: what she will say before (or as) she speaks, depending on latency budget.
Is the technology too early?
| Capability | 2026 readiness for stage |
|---|---|
| Branded character + voice we already love | Ready (Lisa pack) |
| Scripted welcome / intro / sponsor read | Ready if vault-locked |
| Called-on interaction with host (“How are you?” + short bit) | Ready with rehearsal — expect 1–3s lag; write bits that tolerate it |
| Long free-form co-hosting, comedy, crisis MC | Early / risky — don’t bet the show |
| Perfect venue networking | Always the weak link — plan offline fallbacks |
Bottom line: working toward “Hi Lisa” is exactly the right ambition for CBS. Shipping a reliable segment co-host is on the table; shipping a full AI co-MC is not required to innovate or wow the room.
Is it too much to ask?
| Scope | Verdict |
|---|---|
| Full autonomous MC for multi-day agenda, unscripted Q&A, jokes, crisis control | Too much for a first year — latency, hallucination, brand risk |
| Called-on co-host: Mike triggers segments; avatar speaks prepared or tightly grounded replies; human always can cut | Realistic innovator — same energy as CBS innovation, controlled blast radius |
| Segment specialist: only intros, only session summaries, only sponsor reads, only “what just happened” recaps | Best first ship — demo wow + recoverable failures |
Recommendation: ship as a named persona (strong default: Lisa Yuson, already character-complete) inside a reusable stage-agent web runtime, with a visible kill switch, a stage operator, and a shortlist of moments — not free roam.
What “agent co-host” means on stage
Three products must work together in under ~2–4 seconds perceived latency when possible:
1. Face (avatar)
Consistent character image: lip-sync, optional gestures, big-screen safe. Options: Anam.ai, HeyGen live avatars, custom Unreal/Unity, or simpler “voice + fixed portrait / looping video” for v1.
2. Voice
Fixed branded voice (clone or designed). Low latency TTS / streaming. See ElevenLabs, VibeVoice, voice agent notes.
3. Brain
Agent with CBS knowledge vault: agenda, speaker bios, sponsor copy, house rules, “never say” list. Harness: OpenClaw / Hermes / Grok Build / company stack — not raw ChatGPT on Wi‑Fi alone.
Mike (or stage manager) is the fourth layer: human gate — push-to-talk, segment select, interrupt, or “read this script only.”
Jobs the co-host can own (CBS-shaped)
| Moment | Agent role | Risk |
|---|---|---|
| Open / close day | Welcome, house rules, “who’s in the room,” safety/exit notes | Low if scripted |
| Speaker intros | Bio from vault, 30–45s, correct titles, no improvised resume | Med if bio wrong |
| Session summarize | 3 bullets after a talk (notes + transcript feed) | Med — need good capture |
| Sponsor / partner reads | Exact approved copy only | Low if locked text |
| Audience Q filter | Cluster questions; surface top 3 for Mike | Low (assist only) |
| Logistics (“lunch is…”, “track B is…”) | From live ops board / HEOS-style TODAY for the event | Low if ops feed is clean |
| Banter with Mike | Short call-and-response, rehearsed bits | High if free-form |
| Live debate / unscripted expert answers | Defer to humans | Too high for v1 |
Architecture (discussion draft)
┌──────────────────────────────────────────────────────┐
│ STAGE │
│ Mike (mic) ←→ Operator tablet ←→ Kill switch │
│ ↓ │
│ Confidence monitor (what agent will say next) │
│ ↓ │
│ Big screen: Avatar video stream │
│ PA: Voice TTS │
└──────────────────────────────────────────────────────┘
│ Wi‑Fi / hardline preferred
▼
┌──────────────────────────────────────────────────────┐
│ CONTROL PLANE (backstage laptop / mini) │
│ • Segment buttons: Intro | Summarize | Sponsor | Ops │
│ • Script lock vs “open question” mode │
│ • Latency / fallback audio file library │
└──────────────────────────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────────┐
│ BRAIN │
│ CBS vault: agenda, bios, sponsors, rulings │
│ Agent harness (OpenClaw / Hermes / Grok / QM-style) │
│ Optional: multiplayer room for ops team │
│ (GFAVIP Webchat / Slack) │
└──────────────────────────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────────┐
│ RENDER │
│ Avatar provider (Anam / HeyGen / simpler v1) │
│ TTS provider (ElevenLabs / etc.) │
│ Optional STT of Mike for “call by name” │
└──────────────────────────────────────────────────────┘ Same layers as our multi-agent harness map: surface (stage), harness (brain), vault (CBS knowledge) — plus a broadcast renderer.
Human control model (non-negotiable)
- Mike calls the co-host — not continuous free talk unless rehearsed.
- Operator confirms next utterance on a confidence screen (or green-lights segment).
- Hard kill — mute TTS + freeze/hide avatar in <1s (hardware switch if possible).
- Fallback packs — pre-rendered audio/video for intros and sponsors if network dies.
- Rulings file — words never said, legal, sponsor constraints (same idea as knowledge base rulings).
- Permissions — no live send of emails/posts from stage agent without offline review.
Stack options (research shortlist)
| Layer | Options we already document / use | Notes for stage |
|---|---|---|
| Avatar | Anam + OpenClaw, HeyGen, Hyperframes-style video-as-code later | Prefer providers with streaming low-latency; stress-test on venue Wi‑Fi |
| Voice | ElevenLabs, voice agents, VibeVoice | Clone or design a CBS character, not Mike’s exact voice (confusion / uncanny) |
| Brain | OpenClaw, Hermes, Grok Build, HEOS vault | Offline-capable brief + online model; never only cloud chat in browser |
| Ops surface | GFAVIP Webchat, multiplayer room, stage tablet | Ops posts agenda changes; agent reads same source of truth |
| Capture for summaries | Room mics → STT → agent summarize | Consent + recording policy for speakers/attendees |
v0 (lowest risk): fixed character art + excellent TTS + teleprompter-style agent-generated script on operator screen that Mike or co-host reads — still “AI assisted,” less avatar risk. v1: live avatar + locked segments. v2: open Q with heavy grounding after v1 works twice in dress rehearsal.
CBS knowledge vault (what the agent must know)
Treat as a mini HEOS / knowledge base for the event:
agenda.md— day, room, times (single source; ops owns updates)speakers/*.md— approved bios onlysponsors/*.md— approved reads, logos, never-claimsvoice.md— co-host tone (warm, cross-border, not salesy, not political)rulings.md— never improvise legal, medical, investment advice; defer to Mikerun-of-show.md— when avatar is allowed on, cues, music beds
Fill with reverse-prompting style interviews of the production team (Alex Finn reverse prompting) so the vault is complete before tech build.
Risks & mitigations
| Risk | Mitigation |
|---|---|
| Venue Wi‑Fi dies | Hardline or local hotspot; offline fallback media; LTE backup |
| Latency kills banter | Segment-based not free conversation; pre-generate where possible |
| Wrong bio / sponsor claim | Script lock from vault; operator preview; no free web search on stage |
| Uncanny / creepy brand hit | Friendly character design; short segments; humor that doesn’t mock speakers |
| Recording / likeness rights | Don’t clone a real person without contracts; original avatar preferred |
| Demo becomes the story (failure) | Position as experiment; Mike always can finish alone |
Path to November (rough timeline)
- Now (discussion): this page — agree scope (v0/v1), persona brief, kill-switch owner.
- +2–3 weeks: vault draft (agenda shape, sample bios, voice/rulings); pick avatar+voice vendor for a lab demo.
- +6 weeks: closed-room demo: Mike calls 3 segments (intro, summary, sponsor) with operator gate.
- +10 weeks: dress rehearsal in similar room (PA, projection, latency); fallback pack ready.
- Show week: limited live moments only; record for GTM/content after.
Parallel: CBS GTM and attendee systems stay human-led — GTM plan, GFAVIP perks — co-host is production innovation, not a substitute for ops.
Open decisions (for the team)
- Primary persona: Lisa Yuson vs dedicated CBS-only character vs multi-roster later?
- Product codename for the reusable app (
stage-agent/ Showhost / other)? - Brain: OpenClaw (Lisa continuity) vs separate low-latency stage brain with same soul files?
- v0 (voice + still/looping portrait) vs v1 (live lip-sync avatar) for first public use?
- Who is stage operator / kill-switch owner?
- Which moments are guaranteed live vs pre-rendered backup?
- Recording + release rights for co-host segments (and Lisa brand guidelines)?
- Budget: avatar SaaS + voice + on-site hardware + engineer days?
Related on this site
Bottom line
Yes — work toward “Hi Lisa” on a dedicated stage display. The tech for a called co-host is ready enough if we treat this as a reusable stage-agent web runtime (presence + console + vault + kill switch), plug in OpenClaw personas like Lisa as packs, and bound the show to rehearsed segments. Not too early to innovate; too early to let the agent run the conference unsupervised.
Product brief: Stage Agent Runtime
· Agent build handoff: procedures/stage-agent-build-brief.md
(pass into Grok Build /goal — Phase 0 scaffold first).
Comments
Approved comments appear below. Log in once with GFAVIP — it applies across the whole site. GFAVIP login
View comments archive