Skip to content

Differentiable Simulation

NovaPhy's differentiable path is being fused back into the regular physics data model. New differentiable work should start from the same Model and SimState objects used by the forward engine, then expose the required DeviceArray storage to Nova as zero-copy tensor views.

The old tensor-only compatibility directories collision_nova and sim_nova have been removed. The supported path uses regular engine objects plus zero-copy Nova views.

All numerical arrays in the differentiable Python runtime and demos use numpy.float32, matching the engine's C++ float and Nova F32 tensor contract. Diagnostic and finite-difference calculations do not introduce a separate float64 path.

Forward and Gradient Contract

DiffSemiImplicitSolver.step is a forward state transition. It exists because reverse-mode AD must first execute and record the operations whose adjoint will later be replayed by Tape.backward; backward does not replace the forward step. The solver accepts regular Model, SimState, Control, and Contacts objects, and its Tape differentiates the generated operations that this same step actually executed.

This does not mean that every production solver already has an exact VJP. DiffSemiImplicitSolver currently covers a subset of semi-implicit particle and rigid dynamics. It is not the backward implementation of SolverSemiImplicit, VBD, or the complete CollisionPipeline. Standalone generated rollouts used by the soft-body and dice demos are smooth optimization surrogates; their gradients are gradients of those generated rollouts, not of an unrecorded production VBD/contact step. The guarded planar 2R recorder is exact only in its documented domain. General Featherstone sensitivity is a frozen-dynamics linearization and is named and documented as such.

Build Switches

Flag Default Purpose
NOVAPHY_WITH_DIFF OFF Build Nova tensor/autodiff support and fused differentiable physics kernels explicitly.
NOVAPHY_WITH_CUDA OFF Build the core CUDA DeviceArray backend. Required when a differentiable demo finalizes a regular Model on CUDA.
NOVAPHY_BUILD_VBD ON Build the VBD/AVBD module and Python bindings. Set this to OFF to remove VBD from the build.
NOVAPHY_WITH_VBD_CUDA OFF Build the VBD CUDA backend. Requires NOVAPHY_BUILD_VBD=ON.
NOVAPHY_WITH_VBD_DLAN OFF Build the VBD Denglin backend. Requires NOVAPHY_BUILD_VBD=ON.
NOVAPHY_WITH_COLLISION_CUDA ON Build the shared non-differentiable CUDA collision pipeline. Required by the generic diff-ball path on CUDA.
NOVA_ENABLE_NVRTC ON Build the NVRTC JIT backend for generated Nova kernels. Required for the fast fused-step path.

For the current CUDA diff-ball development build:

conda run -n novaphy_diff env \
  CMAKE_BUILD_PARALLEL_LEVEL=1 \
  CMAKE_ARGS="-DCMAKE_PREFIX_PATH=/path/to/NovaPhy/build/vcpkg_installed/x64-linux;/path/to/conda/env/lib/python3.11/site-packages \
  -Dpybind11_DIR=/path/to/conda/env/lib/python3.11/site-packages/pybind11/share/cmake/pybind11 \
  -DNOVAPHY_BUILD_TESTS=OFF \
  -DNOVAPHY_BUILD_VBD=OFF \
  -DNOVAPHY_WITH_CUDA=ON \
  -DNOVAPHY_WITH_COLLISION_CUDA=ON \
  -DNOVA_ENABLE_NVRTC=ON \
  -DNOVAPHY_CUDA_ARCHITECTURES=86 \
  -DCMAKE_CUDA_ARCHITECTURES=86" \
  python -m pip install -e . --no-build-isolation

Use the architecture value for the local GPU. Passing it explicitly avoids CUDA 12 toolchains accidentally selecting unsupported native architectures on newer cards.

Runtime Layers

The fused path has three layers:

Layer Files Role
Zero-copy views include/diff/tensor_views.h, src/diff/tensor_views.cpp Wrap selected Model and SimState DeviceArray buffers as nova::Tensor external views. This keeps layout conversion out of physics kernels.
Differentiable collision forces include/collision/diff_soft_contacts.h, src/collision/diff_soft_contacts.cpp Generated indexed Nova kernels that consume regular Contacts.soft_contact_* buffers. Current differentiable coverage is particle contact against world-owned static shapes (shape_body < 0); body-attached static or kinematic shapes are outside this operator's contract.
Differentiable particle dynamics include/dynamics/semi_implicit/diff_particles.h, src/dynamics/semi_implicit/diff_particles.cpp Generated indexed Nova kernels for particle semi-implicit integration and point loss.
Differentiable rigid dynamics include/dynamics/semi_implicit/diff_rigid.h, src/dynamics/semi_implicit/diff_rigid.cpp Generated indexed Nova kernels for static box-plane contact accumulation, rigid-body semi-implicit integration, and rigid terminal losses.
Shared traced contact formula include/collision/diff_soft_contact_trace.h Single source of truth for static particle-shape contact math used by force-only and fused-step kernels.

Non-differentiable collision broad phase, topology, and contact bookkeeping should continue to use src/collision/collision_pipeline.cpp. Differentiable force accumulation consumes the resulting regular Contacts and the same Model/SimState storage through tensor views rather than copying into a second simulation model.

Python API

Import Nova runtime and fused differentiable helpers with:

import novaphy
from novaphy import nova
import novaphy.diff as ndiff
from novaphy.solvers import DiffSemiImplicitSolver, DiffSolverConfig

novaphy.solvers is the canonical solver namespace. Generated operations and tensor-view helpers remain under novaphy.diff; differentiable solver types are intentionally not re-exported from the top-level novaphy namespace.

Tensor Views

API Purpose
ndiff.state_particle_q_tensor(state, requires_grad=False) Zero-copy view of SimState.particle_q.
ndiff.state_particle_qd_tensor(state, requires_grad=False) Zero-copy view of SimState.particle_qd.
ndiff.state_particle_f_tensor(state, requires_grad=False) Zero-copy view of SimState.particle_f.
ndiff.state_body_q_tensor(state, requires_grad=False) Zero-copy {body_count, 7} view of SimState.body_q.
ndiff.state_body_qd_tensor(state, requires_grad=False) Zero-copy {body_count, 6} view of SimState.body_qd.
ndiff.state_body_f_tensor(state, requires_grad=False) Zero-copy {body_count, 6} view of SimState.body_f.
ndiff.state_joint_q_tensor(state, requires_grad=False) Zero-copy view of SimState.joint_q.
ndiff.state_joint_qd_tensor(state, requires_grad=False) Zero-copy view of SimState.joint_qd.
ndiff.model_tet_activation_tensor(model, requires_grad=False) Zero-copy view of Model.particle_tet_activation.
ndiff.model_tet_material_tensor(model, requires_grad=False) Zero-copy view of Model.particle_tet_material.
ndiff.model_tri_activation_tensor(model, requires_grad=False) Zero-copy view of Model.particle_tri_activation.
ndiff.model_tri_material_tensor(model, requires_grad=False) Zero-copy view of Model.particle_tri_material.

These views share storage with the original SimState. Set requires_grad=True only for data that should receive adjoints.

Generated Physics Kernels

API Purpose
ndiff.eval_static_particle_shape_contact_forces_generated(model, particle_q, particle_qd, particle_f, soft_contact_margin=0.0) Returns particle forces after world-owned static box/plane soft contacts are accumulated into particle_f.
ndiff.assign_tet_materials_generated(material_params, base_tet_materials, tet_count) Builds per-tet material tensors from either 2 shared material parameters or 2 parameters per tet, preserving damping from the base material tensor.
ndiff.eval_tetrahedra_forces_generated(model, particle_q, particle_qd, particle_f, tet_activations, tet_materials) Returns particle forces after Newton-style tetrahedral FEM forces are accumulated into particle_f.
ndiff.eval_triangle_forces_generated(model, particle_q, particle_qd, particle_f, tri_activations, tri_materials) Returns particle forces after Newton-style triangle membrane/area/drag/lift forces are accumulated into particle_f.
ndiff.integrate_particles_generated(model, particle_q, particle_qd, particle_f, dt) Returns (particle_q_next, particle_qd_next) using semi-implicit particle integration.
ndiff.integrate_particles_with_static_contacts_generated(model, particle_q, particle_qd, particle_f, soft_contact_margin, dt) Returns (particle_q_next, particle_qd_next) after accumulating world-owned static box/plane soft contacts and integrating particles inside one generated indexed kernel.
ndiff.integrate_rigid_bodies_with_static_contacts_generated(model, body_q, body_qd, body_f, angular_damping, friction_smoothing, dt) Returns (body_q_next, body_qd_next) after static box-plane penalty contacts and rigid semi-implicit integration in one generated indexed kernel.
ndiff.body_twist_from_omega_generated(omega, linear_x, linear_y, linear_z) Builds packed rigid twists from differentiable angular velocity and fixed linear velocity.
ndiff.rigid_orientation_velocity_loss_generated(...) Generic two-axis orientation and optional height loss with terminal linear/angular velocity regularization.
ndiff.particle_point_loss_generated(particle_q, particle_index, target_x, target_y, target_z) Returns a scalar squared-distance loss for one particle.
ndiff.particle_centroid_generated(particle_q) Returns the unweighted {1, 3} centroid of the particle positions.
ndiff.vec3_point_loss_generated(point, target_x, target_y, target_z) Returns a scalar squared-distance loss from a vec3 tensor to a target.
ndiff.clamp_f32_generated(x, lower, upper) Returns a graph-capturable clamped tensor for projected parameter updates.
ndiff.record_planar_2r_featherstone_step(model, state_in, control, q, qd, q_out, qd_out, dt, angular_damping) Records the guarded analytic VJP for one supported, contact-free planar 2R SolverFeatherstone step.
ndiff.planar_2r_tip_loss(...) Analytic end-effector squared-distance loss for the planar 2R articulation.
ndiff.integrate_push_box_generated(...) Differentiable one-axis sphere-pusher/box penalty contact and box translation step.

The kernels use nova::make_indexed_expr_kernel_typed_access_n and nova::make_indexed_fixed_reduce_expr_kernel_typed_access_n, so model arrays such as radii, flags, transforms, material parameters, and shape metadata are read as source tensors. Dynamic state tensors such as particle_q and particle_qd can require gradients.

Soft-body rest triangles and tetrahedra must be nondegenerate; ModelBuilder rejects near-zero rest area/volume and inverted tetrahedron orientation. Generated FEM treats zero volumetric stiffness as its finite zero-force limit instead of evaluating a 0 / 0 material ratio.

Generated particle/contact/FEM vector rows are tightly packed {n, 3}; material rows likewise use their documented exact widths. Wider matrices are rejected because these generated expressions use fixed scalarized row widths, not arbitrary tensor strides.

Capture Graph Pattern

nova.ScopedCapture() records Nova tensor/kernel launches whose storage and launch sequence stay fixed. It does not record the DeviceArray clear, copy, snapshot, host staging, or component-mirror operations performed by a complete DiffSemiImplicitSolver.step(). Consequently that solver reports supports_graph_capture=False and must not be wrapped in Nova logical capture.

Pure Nova operator sequences, including tape.backward(), tensor gradient clearing, and optimizer kernels, may still use ScopedCapture when all their inputs and outputs are fixed Nova tensors. Driver-layer capture remains the authoritative mechanism for engine operations that explicitly support it.

For vec3 parameters that need the same finite-gradient and norm clipping used by Newton's dice example:

nova.sgd_update_vec3_clipped_f32(
    omega, learning_rate, max_parameter_norm, max_gradient_norm
)

For performance, enable Nova JIT before constructing generated kernels:

nova.set_jit_enabled(True)

The diff demos do this by default. If Nova was built without NOVA_ENABLE_NVRTC, the same code falls back to the bytecode interpreter.

Diffsim Ball Demo

python/demos/diff/demo_diffsim_ball.py follows Newton's example_diffsim_ball.py framework path:

  • The scene is built through the regular ModelBuilder.
  • The ball is a regular particle in a regular Model/SimState.
  • The floor is a regular static Z-up plane shape and the vertical wall is a regular static box shape.
  • CollisionPipeline.collide(states[0], contacts) runs once outside the Tape, matching Newton's one-shot static wall/ground contact generation.
  • Every substep calls states[t].clear_forces() and DiffSemiImplicitSolver.step(states[t], states[t + 1], control, contacts, dt).
  • The solver consumes Contacts.soft_contact_*, evaluates contact force in one generated kernel, and integrates particles in a second generated kernel.
  • Backward uses the generic Nova tape and generated adjoints, not a demo-specific backward path.
  • The GUI trajectory is refreshed every few training iterations, so the drawn path follows the optimized ball trajectory.

Run a headless smoke test:

conda run -n novaphy_diff python python/demos/diff/demo_diffsim_ball.py --headless 5

Expected behavior is a decreasing loss and changing optimized initial velocity.

The initial 289-substep trajectory and final velocity compare bit-for-bit with Newton on the current CUDA development build. AD versus central finite difference has a maximum relative error of about 2.2e-3.

Use --dump-states path.npz with --headless 0 to write the initial rollout for Newton/NovaPhy comparison. Passing --headless N first performs N training iterations, then dumps the optimized trajectory. The NPZ contains frame-level particle_q shaped {sim_steps + 1, particle_count, 3} and also particle_q_substeps shaped {sim_steps * sim_substeps + 1, particle_count, 3}.

Diffsim Soft Body Demo

python/demos/diff/demo_diffsim_soft_body.py is a generated soft-body material-optimization surrogate matching Newton's example_diffsim_soft_body.py workflow:

  • The soft grid, Z-up ground plane, and wall are built through the regular ModelBuilder.
  • The optimized material parameters are stored as a Nova tensor and expanded to per-tet material tensors by assign_tet_materials_generated.
  • Each substep accumulates triangle forces, tetrahedral FEM forces, static particle-shape soft contact forces, and then integrates particles.
  • The loss is the squared distance between particle center of mass and the target.
  • Backward uses the generic Nova tape and generated adjoints.

The model and state storage are shared with the regular engine, but this rollout does not call production VBD. Its gradient therefore belongs to the generated FEM/contact/integration sequence listed above.

Run a headless smoke test:

conda run -n novaphy_diff python python/demos/diff/demo_diffsim_soft_body.py --headless 3

Use --dump-states path.npz --headless 0 to write the initial rollout for Newton/NovaPhy comparison. Passing --headless N first performs N training iterations, then dumps the optimized trajectory.

Diffsim Dice Demo

python/demos/diff/demo_diffsim_dice.py is a smooth generated surrogate of Newton's example_diffsim_dice.py rigid optimization workflow:

  • The dice and Z-up ground are regular ModelBuilder shapes.
  • Candidate shape pairs come from the regular model's shape_contact_pairs, the same explicit broad-phase topology consumed by CollisionPipeline.
  • Newton's mass construction is preserved: the body starts with the requested 4 kg base mass and the box contributes its default 1000 kg/m3 density.
  • Static plane-box narrow phase keeps the four deepest box corners, applies shape margins to the contact points, and uses the same averaged penalty material, Huber-smoothed friction, gyroscopic term, and semi-implicit rigid integration.
  • The generated rigid integrator uses the full body-frame inverse inertia, including world/body coordinate transforms and the gyroscopic term.
  • One generated kernel performs all contact force/torque accumulation and integration for one substep. Nova generates its adjoint; there is no dice-specific backward implementation.
  • Forward/backward kernels run as ordinary Nova launches. The host-side line search evaluates candidate trajectories with the regular forward rollout; this demo does not place the full training step in one CUDA graph.
  • ViewerGL replays each refreshed trajectory frame by frame and retains prior training trajectories.

This demo reuses production model data and broad-phase topology, but it does not differentiate a CollisionPipeline plus SolverSemiImplicit call. Its gradient is the derivative of the generated box-plane penalty/contact rollout.

Run:

conda run -n novaphy_diff python python/demos/diff/demo_diffsim_dice.py
conda run -n novaphy_diff python python/demos/diff/demo_diffsim_dice.py --headless 10
conda run -n novaphy_diff python python/demos/diff/demo_diffsim_dice.py --check-grad

The contact-gradient check uses a deliberately penetrating box-plane state away from the discrete contact-generation threshold.

Diffsim Featherstone Demo

python/demos/diff/demo_diffsim_featherstone.py ports Newton's example_diffsim_feathstone.py initial-hinge-velocity optimization:

  • the forward rollout is the regular CUDA SolverFeatherstone operating on ordinary Model and SimState objects;
  • joint_q and joint_qd enter Nova through cached zero-copy tensor views;
  • every contact-free step records a hand-derived analytic VJP for the two-link planar revolute chain (mass matrix, Coriolis and gravity terms, joint limits, semi-implicit integration, and angular damping);
  • the viewer keeps the arm in its vertical XOZ plane, colors the links blue and orange, renders a black-and-white checkerboard horizontal ground, and draws the end-effector trajectory in red;
  • no collision_nova or sim_nova compatibility object is used.

Run the optimization or the full 576-substep gradient check with:

conda run -n novaphy_diff python python/demos/diff/demo_diffsim_featherstone.py --headless 20
conda run -n novaphy_diff python python/demos/diff/demo_diffsim_featherstone.py --check-grad --headless 0

The analytic recorder rejects unsupported topology and solver features explicitly. Its supported domain is one contact-free, world-rooted two-link articulation with two enabled +Y revolute joints, dynamic bodies, identity joint rotations, link COMs at their body origins, zero target drives, friction, joint damping, feedforward force, and external body force. It also checks q_out / qd_out against the analytic update before recording the VJP, so an active effort/velocity clamp, contact, or other unmodeled forward effect fails instead of silently returning a surrogate gradient.

Differentiable Ant Policy Fine-Tuning

python/demos/diff/demo_diffsim_ant_walk.py combines the D.VA Ant morphology, a pretrained SB3 SAC Ant-v3 actor, the regular Featherstone/PGS forward solver, and an analytic direct-force sensitivity. The default playback remains a pure forward-policy demo. Passing --train N fine-tunes the eight biases in the actor's final pre-tanh layer before replay.

The training gradient crosses every 1 ms Featherstone substep in the selected multi-frame horizon. It uses generalized mass matrices, spatial Jacobians, and a projection into the active contact-normal null space. Contact creation, state-dependent mass/bias derivatives, friction tangents, and gradients across the 20 Hz observation refresh are frozen. A real forward line search guards this approximation: a candidate is retained only if the ordinary solver measures a lower trajectory loss.

conda run -n novaphy_diff python \
  python/demos/diff/demo_diffsim_ant_walk.py --check-gradient

conda run -n novaphy_diff python \
  python/demos/diff/demo_diffsim_ant_walk.py \
  --train 10 --train-horizon 8 \
  --save-policy /tmp/novaphy_ant_trained.npz --headless 200

conda run -n novaphy_diff python \
  python/demos/diff/demo_diffsim_ant_walk.py \
  --policy /tmp/novaphy_ant_trained.npz

From-zero periodic gait

python/demos/diff/demo_diffsim_ant_from_scratch.py removes the pretrained SAC policy. Its 80-parameter Actor starts at exact zeros, so the initial rollout applies no joint torque. A horizon curriculum first optimizes standing and 4, 8, 12, 16, 24, 36, and 60 policy-frame trajectories, followed by a 400-frame stabilization stage. The Actor combines first/second-harmonic phase features with torso height, velocity, and tilt feedback. Its objective additionally penalizes torso tilt/heading error, lateral drift, joint-pose distortion, action discontinuity, and violations of diagonal-leg symmetry.

# Train from zero and open ViewerGL when training finishes.
conda run -n novaphy_diff python \
  python/demos/diff/demo_diffsim_ant_from_scratch.py

# Replay the checked-in result without training.
conda run -n novaphy_diff python \
  python/demos/diff/demo_diffsim_ant_from_scratch.py \
  --load python/demos/data/ant/rl_policies/periodic_ant_natural_gait.npz

The saved natural-gait run remains upright for a 600-frame/30-second validation, travels 0.792 m, ends with positive forward velocity, and limits maximum lateral drift to 0.253 m and torso tilt to 5.32 degrees. This is still a slow nominal-terrain gait rather than a robust locomotion policy: no disturbance or terrain randomization is included.

Hybrid Franka Push-Box Demo

python/demos/diff/demo_diffsim_franka_push_box.py demonstrates a hybrid gradient boundary on the repository's FR3 asset:

  • regular CPU SolverFeatherstone advances the seven-axis arm;
  • novaphy.solvers.FeatherstoneAnalyticControlSensitivity propagates an arbitrary controller parameterization through the generalized mass matrix, implicit PD response, and spatial point Jacobian;
  • the end-effector position/velocity trajectory is passed to integrate_push_box_generated, whose smooth penalty contact and one-axis box integration are differentiated by Nova Tape;
  • the contact trajectory adjoints are multiplied by the analytic Featherstone trajectory Jacobian to obtain the key-pose gradient.

The reusable control-sensitivity backend supports a selected articulation containing fixed and one-DOF joints. It is a frozen-dynamics linearization: collision active-set changes and derivatives of state-dependent mass/bias terms are not yet a complete general SolverFeatherstone.step_vjp. The demo also assumes a stiff position-controlled pusher, so the generated contact operator moves the box but does not feed its reaction force back into the arm.

Run the GUI, headless training, or directional gradient check with:

conda run -n novaphy_diff python \
  python/demos/diff/demo_diffsim_franka_push_box.py
conda run -n novaphy_diff python \
  python/demos/diff/demo_diffsim_franka_push_box.py --train 5
conda run -n novaphy_diff python \
  python/demos/diff/demo_diffsim_franka_push_box.py --check-gradient

Current Coverage

Area Status
Zero-copy SimState particle position/velocity/force views Implemented.
Zero-copy SimState rigid pose/twist/force views Implemented.
Zero-copy Model particle and static-shape source views Implemented internally.
World-owned static particle-vs-box soft contact Differentiable generated path implemented.
World-owned static particle-vs-plane soft contact Differentiable generated path implemented.
One-axis sphere-pusher/box contact Differentiable generated path implemented for the hybrid Franka demo.
Generic Featherstone PD control sensitivity Analytic frozen-dynamics backend implemented for selected fixed/one-DOF articulations.
Fused particle static-contact integration Differentiable generated path implemented with fixed source-domain reduction and NVRTC JIT.
Tetrahedral FEM force Differentiable generated path implemented for the Newton eval_tetrahedra formula.
Triangle membrane/area/drag/lift force Differentiable generated path implemented for the Newton eval_triangle formula.
Soft-body material assignment Differentiable generated path implemented for shared or per-tet mu/lambda parameters.
Particle semi-implicit integration Differentiable generated path implemented.
Point target loss Differentiable generated path implemented.
Particle center-of-mass loss building blocks Differentiable generated path implemented.
Dynamic box-vs-static-plane rigid contact Differentiable generated path implemented, including four-contact reduction and semi-implicit integration.
General dynamic rigid body contacts Not fused yet. Use the shared non-differentiable collision pipeline for unsupported topology/bookkeeping.
Mesh contacts Not fused yet.
Spring and bending forces Not fused yet.
Soft-body static primitive shape contacts Differentiable generated path implemented for particle-vs-box/plane.
Soft-body dynamic/mesh contact topology Not fused yet; collision topology still needs to come from the shared collision pipeline.
VBD Optional module; can be fully disabled with NOVAPHY_BUILD_VBD=OFF.

Design Rules

  • Keep physics logic shared with the forward engine wherever possible.
  • Keep zero-copy adaptation in diff/tensor_views.*.
  • Put differentiable math in domain files beside the forward code, such as collision/diff_soft_contacts.* and dynamics/semi_implicit/diff_particles.*.
  • Prefer generated indexed Nova kernels for differentiable forward and adjoint code.
  • Use collision_pipeline.cpp for non-differentiable broad phase and contact topology instead of creating a second collision stack.
  • Do not add new dependencies on collision_nova or sim_nova.