Authors

Document Type

Theses, Ph.D

Abstract

Virtual characters require animation capable of portraying dynamic, context-sensitive human-like behaviours. Several approaches to generating such animation have been developed, but each carries limitations. Motion capture can produce high-fidelity animation but is expensive and ill-suited to systems that must respond in real time. Physics-based reinforcement learning (RL) enables flexible, dynamic behaviour portrayal, yet relies on simulation feedback signals that are unavailable for social gestures. Supervised approaches can learn social behaviours from motion capture data but yield agents with limited flexibility and generalisation.

This thesis presents RLAnimate, a model-based, data-driven RL framework for character animation that enables a single agent to portray multiple types of interactive human-like behaviours through uniform input parameterisation. The research addresses four progressive research questions, culminating in beat gesture generation from speech(RQ4), with earlier questions establishing reinforcement learning viability (RQ1), human-like quality (RQ2), and fine-grained animation control (RQ3). RLAnimate agents incorporate latent state space models that capture animation dynamics, learning how joint rotation sequences produce human-like motion. Compact motion capture datasets inform the training process, requiring significantly less data than su pervised alternatives. The framework is first shown to produce dynamic, human-like waving and pointing behaviours across a wide range of intensities, targets, and durations.

To address the further challenges of portraying complex, speech-driven beat gestures, the framework is extended with an augmented animation generation technique and a realism regularisation regimen. A perceptual evaluation comparing RLAnimate to an existing method for beat gesture portrayal found that agents trained with realism regularisation are statistically indistinguishable from human motion. In its current form, RLAnimate addresses upper-body animation only. Despite this limitation, the framework demonstrates that model-based RL with uniform parameterisation can produce dynamic, frame-by-frame portrayal of multiple human-like behaviours, establishing a viable alternative to both physics-based and supervised approaches for social gesture synthesis.

DOI

https://doi.org/10.21427/aw1h-3j07

Creative Commons License

Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License
This work is licensed under a Creative Commons Attribution-NonCommercial-Share Alike 4.0 International License.


Share

COinS