Rig-Your-Portrait: Controllable and Relightable Portrait Video Generation with Explicit 3D Guidance

Dingyun Zhang

University of Science and Technology of China


🌿 In April 2024, I proposed and implemented a diffusion-based portrait animation method. The method can be simply described as identity-preserving controllable and relightable portrait video generation based on 3D explicit guidance. Trained on a limited dataset in my spare time, I named this method Rig-Your-Portrait. At the time, I decided to qualitatively present this method in my thesis, and this has been completed. I was also doing the internship and, for the purpose of communication, presented this method along with preliminary results to the group members in April. This webpage introduces the methodology and the preliminary results of the project.

Inspired by Animate Anyone and DiffusionRig, a tuning-based method for facial image editing, Rig-Your-Portrait extends the Animate Anyone paradigm to controllable portrait video generation, adapting the network to the specific characteristics of the portrait.

Limitations of existing methods: existing diffusion-based portrait animation methods are generally capable of controlling the expressions and head poses of a reference portrait using driving frames, but lack illumination control. Furthermore, the control signals in these methods are primarily based on 2D explicit guidance, such as the 2D facial landmarks used in AniPortrait. During training, models are typically trained on image pairs of the same identity. As a result, during cross-identity animation in the inference stage, misalignment between the 2D driving signals and the corresponding features of the reference portrait leads to discrepancies in facial shape between the animated image and the reference image, i.e., identity leakage issue.

The goal of Rig-Your-Portrait: given a reference portrait and a driving video sequence (a portrait or a video), the model can disentangledly transfer the facial expressions, head poses, or illumination attributes from the driving frames to the reference portrait, enabling controllable portrait video generation.

Specifically, in cross-reenactment, by leveraging the facial attribute disentanglement characteristics of the explicit 3D face reconstruction models, Rig-Your-Portrait aims to preserve the identity-related information (e.g., facial shape) of the reference portrait, thereby alleviating the identity leakage issues (such as unintended facial shape alterations) in the generated videos.


☘️ Method

Method overview of Rig-Your-Portrait. Left side: the training pipeline. Right side: the inference pipeline.


🚗 Model Architecture and Parameter Initialization

The model network comprises a VAE (including encoder and decoder, i.e., VAE enc. and VAE dec.), a CLIP image encoder, 3D face reconstruction models, DenoisingNet, ReferenceNet, an expression guidance encoder (Exp enc.), a geometry guidance encoder (Geo enc.), a texture guidance encoder (Tex enc.), and a Lambertian illumination guidance encoder (Lamb enc.).

The architecture and weights of the VAE, DenoisingNet, and ReferenceNet are inherited from Stable Diffusion, while the temporal layer of DenoisingNet is initialized with the pretrained weights from AnimateDiff. The 3D face reconstruction models are SMIRK and DECA, used during the data processing stage. Additionally, the expression, geometry, texture, and illumination guidance encoders share the same lightweight architecture, utilizing four convolutional layers for feature extraction.


🌲 Training Stage

The model is trained on a small subset of publicly available datasets, including a limited number of portrait videos with different identities, expressions, and head poses, as well as partial indoor multi-illumination images from Multi-PIE. Multi-PIE, which consists of multi-view static illumination portraits captured under indoor white light with limited incident angles, is used to preliminarily validate the feasibility of illumination control. The model's loss function inherits from Animate Anyone, with the training process being conducted in two stages.

In the first training stage, DenoisingNet (excluding the temporal layer), ReferenceNet, and the expression, geometry, texture, and illumination guidance encoders are optimized. The reference image \( I_s \) is randomly selected from the entire portrait video clip or pseudo-video clip (where Multi-PIE images of the same identity and camera pose under varying illumination conditions are treated as a video).

3D explicit guidance maps are generated by extracting FLAME parameters from the video frame \( I_d \) using 3D face reconstruction models:

Model inputs include \( I_s \) and the rendered guidance sequences (expression, geometry, texture, illumination) from \( I_d \).

The 3D explicit guidance maps are processed by their respective encoders (Exp enc., Geo enc., Tex enc., Lamb enc.). The extracted features are summed, then fused with the noisy latent 'Noise', and injected into DenoisingNet.

The model extracts appearance features from the reference image using ReferenceNet, integrating them into DenoisingNet via spatial-attention. Simultaneously, the CLIP image encoder encodes the semantic features of the reference image for cross-attention.

In the second training stage, the temporal layer is incorporated after the spatial-attention and cross-attention components in DenoisingNet. Only the temporal layer is updated, while all other parameters remain frozen.


✨ Inference Stage

During inference, all network weights are frozen. For facial attribute transfer, FLAME parameters are extracted from both the reference image \( I_s \) and driving sequence \( I_d \). When specifying attributes (e.g., expression and head pose), the reference portrait's expression and pose coefficients are replaced with those from the driving sequence. The modified FLAME parameters are rendered into 3D guidance maps, which are sequentially fed into the guidance encoders (expression, geometry, texture, illumination) as driving signals. Specifically, by keeping the identity parameters in the FLAME parameters of the reference portrait unchanged, the animated image can better preserve the facial shape of the reference portrait, thereby alleviating the identity leakage issue.

The first-stage model enables controllable and relightable portrait editing by transferring specified attributes (expression, head pose, or illumination) from the target image to the reference portrait. The second-stage model extends this to portrait video generation by transferring the facial attributes from the driving video to the reference portrait.


🦋 Preliminary Qualitative Results

🖼️ The first-stage model could disentangledly transfer the facial attributes, such as expressions, head poses, and illumination conditions, from the target portrait to the reference portrait.

As shown above, we use the first-stage model to perform facial attribute transfer on test images from the FFHQ dataset, which was not involved in training. The results show that the model can transfer the target portrait's expression, pose, and illumination attributes to the reference portrait—individually or simultaneously—while preserving the reference portrait's identity features such as facial shape and background.

🎞️ Qualitative results of self-reenactment using the second-stage model.

🎞️ Qualitative results of cross-reenactment using the second-stage model.

As shown above, self-reenactment and cross-reenactment experiments demonstrate simultaneous expression and head pose transfer using the second-stage model. The reference images and the driving videos are from the TalkingHead-1KH and VFHQ test sets. Preliminary qualitative results indicate the method's feasibility and potential for improvement on larger training sets.

🖼️ The inference results of the first-stage model when using the original 3D explicit maps of another test portrait as driving signals.

As demonstrated above, we also present the inference results of the first-stage model (trained without Multi-PIE) when using the original 3D explicit maps of another test portrait as driving signals. As anticipated, the facial shape in the generated results differs from that of the reference images.


☘️ BibTeX


If you would like to document the project blog, you could use the following BibTeX:


        @misc{zhang2025rigyourportrait,
            title={Rig-Your-Portrait: Controllable and Relightable Portrait Video Generation with Explicit 3D Guidance},
            author={Dingyun Zhang},
            year={2025},
            howpublished={\url{https://rigyourportrait.github.io}},
        }
    

Page created in Feb 24, 2025.
© 2025 All rights reserved.