Works

Livestream Architecture

Developed an interactive real-time project on the Twitch gaming platform that enabled dynamic audience engagement through live streaming, leveraging NLP tooling to process and respond to user chat interactions.

  • StackUnity, C#, Tensorflow

How it worked

A viewer's chat message moves through a single loop: it's read off Twitch chat, classified by a TensorFlow model into an intent and a room category, and handed to the Unity scene controller, which drops the matching piece of geometry into the shared build.

Twitch chat!vote bedroomChat ingeststream + de-dupeTensorFlow NLPintent + categoryUnity · C#scene controller3D sceneelement updatesraw messagetext batchscene commandclassified intent

Fig. 01 — chat text becomes a rendered scene change in a single pass, with TensorFlow doing the only interpretation step in between.

What the demo shows

A reconstruction of the clip embedded on the original page: a chat rail feeding a live tally, and a cluster of low-poly geometry that grows as votes land.

LIVE · LiveStream Architecture
bedroom  4.8%
living room  3.1%
playground  1.7%
patio  1.05%
viewer_42:!vote patio
modArchitect:!build bedroom
lurker99:nice roofline lol
chatgremlin:!vote playground
yuejin:top spend unlocked
stream_bot:+1 living room

Illustrative reconstruction based on the demo clip, not a frame capture — proportions and copy are approximate.

Six lenses on the loop

Read as a product rather than a demo reel: what job it does for a viewer, and where the design carries weight versus where it's simply implied by the stack.

01 · core loop

Type a command, see a number move, watch geometry appear. No page reload, no account: the entire interaction surface is the chat box a viewer already has open.

02 · users & motivation

Twitch chat is high-volume and low-attention by default. The bet: a visible, shared, cumulative effect (the house grows) turns passive lurking into repeated small actions.

03 · information architecture

A flat set of room categories keeps the vote tally legible at a glance, but caps expressiveness — no sub-choices for style, material, or placement.

04 · interaction model

Voting happens through free-text chat rather than buttons or polls. NLP is the interface, which is expressive but harder to guarantee correct.

05 · feedback signifiers

A percentage readout and new geometry are the only confirmations. There's no per-message acknowledgment, so a viewer can't easily tell if their own message landed.

06 · constraints from the stack

One NLP hop drives the scene directly, so a misclassification becomes a visible, public mistake — there's no moderation or confidence gate in the loop.

Where the design opens up

The tight, single-hop loop is exactly why it's hard to grow: NLP, game logic, and rendering all live in reach of one Unity process.

Problem 1 · coupling

Swapping or retraining the classifier means touching the same codebase that renders the scene — ML iteration speed is bound to game-build iteration speed.

Problem 2 · vocabulary ceiling

A flat category classifier can vote, but can't say "which wall" or "what material." Richer commands need a real grammar, not just a label.

Problem 3 · no safety valve

Nothing sits between "classified" and "rendered," so a misclassified or hostile message becomes a permanent, visible change to a shared scene.

A proposed three-layer extension

Splitting the single hop into a pipeline with a real boundary at each seam: streaming input owns the chat connection and rate control, Python + TensorFlow owns language understanding, and a dedicated C# modeling server owns the architectural rule set and mesh generation.

01 · STREAMING INPUTTwitch IRC / EventSubchat connectionIngest + rate limiterdedupe, per-user cooldownCommand queueordered message streamraw text envelope · {user, msg, ts}02 · PYTHON + TENSORFLOW PROCESSINGTokenizer + normalizerstrip emotes, lowercaseIntent / entity classifierTensorFlow modelGrammar validatorconfidence gateCommand · {action, target, material, confidence}03 · C# MODELING SERVERCommand dispatchersequencing, undo stackArchitectural rule engineresolves valid placementsMesh / module instantiatorbuilds the scene graph diffSceneDelta · WebSocket broadcastUnity render clients + OBS overlay

Fig. 02 — !build bedroom oak crosses three ownership boundaries; each arrow marks a payload-shape change, the seam the original design skipped by fusing NLP output directly to scene state.

Why split it this way: Python owns layer two because the ML ecosystem (TensorFlow, tokenizers, retraining, evaluation) is native there and shouldn't require a Unity rebuild to iterate on. C# owns layer three because rule resolution and mesh instantiation want Unity's own geometry and asset types, and keeping the rule engine outside the classifier means the same validated command can be replayed, tested, or fed from a source other than Twitch chat.

From free text to architectural elements

The proposed grammar keeps the original's low floor — anyone can type !vote bedroom — while giving the classifier enough structure to target a specific element instead of a category tally.

Chat messageParsed intentScene effect
!vote patioVOTE · category=patioVote weight +1, no geometry change
!build patio oakBUILD · category=patio, material=oakOak-textured patio mesh instantiated
!undoUNDOLast module removed, vote restored
!lock roof 5mLOCK · target=roof, duration=5mRoof immutable for 5 min, throttles griefing
nice roofline lolbelow confidence thresholdDropped — no scene effect

What the extra layers buy

→ latency budget

Three network hops instead of one in-process call — target well under half a second, chat-to-render, so the "I typed it, it happened" feeling still holds live.

→ moderation surface

The grammar validator and per-user cooldown become the place to rate-limit spam and gate low-confidence parses — a seam the original single-hop design didn't have.

→ streamer controls

A separate modeling server can expose admin controls (pause, force-undo, reset) without touching the classifier or the Unity render client.

→ testability

Structured commands between layers two and three can be replayed from a fixture file, so the rule engine becomes testable without a live Twitch connection.

→ operational cost

Three deployable services instead of one Unity build — worth it once the classifier needs its own retrain/deploy cycle, not before.

→ what to measure

Command-to-render latency, parse confidence distribution, noise-vs-command ratio, and repeat participation per viewer session.

© 2026 Yue Jin. All Rights Reserved.