Measure
Typed multimodal evidence
Key motion, audio, hand cameras, posture, inertial sensing, pressure, pedals, and gaze are mapped to a specific twin and time scale.
World Models / Human motion / Structured dynamics
A performer-instrument world model for seeing, hearing, and safely improving expert piano skill.
Based on the July 2026 UT Austin - Sony CSL digital-twin formulation. The listener twin is shown in the full concept and deferred in the first implementation.

One performance, one coupled system
Music performance is a chain of causes: skilled motion drives the key and action, the instrument turns touch into sound, and the performer listens and adapts. This project makes that chain queryable as a coupled performer-instrument twin with a controller layer for safe, personalized training.
Measure
Key motion, audio, hand cameras, posture, inertial sensing, pressure, pedals, and gaze are mapped to a specific twin and time scale.
Model
A hand and skill state meets a structured piano action: key, hammer, felt, string, soundboard, and a calibrated room response.
Simulate
Stochastic port-Hamiltonian dynamics provide a physics-aware world model for long-horizon prediction and counterfactual questions.
Improve
A controller can guide a learner through a safe step in skill geometry, then fade assistance as the learner internalizes the coordination.
Visual map
The visual story moves from the physical interfaces to the latent geometry that makes training interventions inspectable.



Workflow
The project is organized as a sequence that keeps the model useful at every stage and makes clear which claims are measured, computed, or still proposed.
Synchronize motion, touch, sound, posture, and control streams.
Recover latent performer and instrument states with structure-aware dynamics.
Roll forward under altered keys, targets, timing, or injected coordination.
Choose a bounded intervention and reduce support as skill becomes controllable.
Technical view
The formulation is broad enough to include a listener, room, and exoskeleton, but the implementation begins with the portion the available corpus can identify.
The performer twin carries hand pose, joint state, muscle activation, skill, and slow adaptation. The instrument twin carries the action, hammer, strings, board, per-piano parameters, and audio readout.
Finger individuation is represented as a low-dimensional behavioral manifold. A Fisher-Rao metric supplies a ruler for a bounded natural-gradient step instead of forcing every player toward one template.
Passive key motion and audio support visualization and selected computation. Stronger claims require hand and skeleton data, physical structure, felt and string models, listening studies, and validation.
Current scope
The long-term target is not a replay machine. It is a coupled simulator that can explain how touch becomes sound, locate a learner in a space of coordination, and design safer paths toward expert performance.