Differentiable Simulation¶
NovaPhy's differentiable path is being fused back into the regular physics
data model. New differentiable work should start from the same Model and
SimState objects used by the forward engine, then expose the required
DeviceArray storage to Nova as zero-copy tensor views.
The old tensor-only compatibility directories collision_nova and sim_nova
have been removed. The supported path uses regular engine objects plus
zero-copy Nova views.
All numerical arrays in the differentiable Python runtime and demos use
numpy.float32, matching the engine's C++ float and Nova F32 tensor
contract. Diagnostic and finite-difference calculations do not introduce a
separate float64 path.
Forward and Gradient Contract¶
DiffSemiImplicitSolver.step is a forward state transition. It exists because
reverse-mode AD must first execute and record the operations whose adjoint will
later be replayed by Tape.backward; backward does not replace the forward
step. The solver accepts regular Model, SimState, Control, and
Contacts objects, and its Tape differentiates the generated operations that
this same step actually executed.
This does not mean that every production solver already has an exact VJP.
DiffSemiImplicitSolver currently covers a subset of semi-implicit particle
and rigid dynamics. It is not the backward implementation of
SolverSemiImplicit, VBD, or the complete CollisionPipeline. Standalone
generated rollouts used by the soft-body and dice demos are smooth optimization
surrogates; their gradients are gradients of those generated rollouts, not of
an unrecorded production VBD/contact step. The guarded planar 2R recorder is
exact only in its documented domain. General Featherstone sensitivity is a
frozen-dynamics linearization and is named and documented as such.
Build Switches¶
| Flag | Default | Purpose |
|---|---|---|
NOVAPHY_WITH_DIFF |
OFF |
Build Nova tensor/autodiff support and fused differentiable physics kernels explicitly. |
NOVAPHY_WITH_CUDA |
OFF |
Build the core CUDA DeviceArray backend. Required when a differentiable demo finalizes a regular Model on CUDA. |
NOVAPHY_BUILD_VBD |
ON |
Build the VBD/AVBD module and Python bindings. Set this to OFF to remove VBD from the build. |
NOVAPHY_WITH_VBD_CUDA |
OFF |
Build the VBD CUDA backend. Requires NOVAPHY_BUILD_VBD=ON. |
NOVAPHY_WITH_VBD_DLAN |
OFF |
Build the VBD Denglin backend. Requires NOVAPHY_BUILD_VBD=ON. |
NOVAPHY_WITH_COLLISION_CUDA |
ON |
Build the shared non-differentiable CUDA collision pipeline. Required by the generic diff-ball path on CUDA. |
NOVA_ENABLE_NVRTC |
ON |
Build the NVRTC JIT backend for generated Nova kernels. Required for the fast fused-step path. |
For the current CUDA diff-ball development build:
conda run -n novaphy_diff env \
CMAKE_BUILD_PARALLEL_LEVEL=1 \
CMAKE_ARGS="-DCMAKE_PREFIX_PATH=/path/to/NovaPhy/build/vcpkg_installed/x64-linux;/path/to/conda/env/lib/python3.11/site-packages \
-Dpybind11_DIR=/path/to/conda/env/lib/python3.11/site-packages/pybind11/share/cmake/pybind11 \
-DNOVAPHY_BUILD_TESTS=OFF \
-DNOVAPHY_BUILD_VBD=OFF \
-DNOVAPHY_WITH_CUDA=ON \
-DNOVAPHY_WITH_COLLISION_CUDA=ON \
-DNOVA_ENABLE_NVRTC=ON \
-DNOVAPHY_CUDA_ARCHITECTURES=86 \
-DCMAKE_CUDA_ARCHITECTURES=86" \
python -m pip install -e . --no-build-isolation
Use the architecture value for the local GPU. Passing it explicitly avoids CUDA 12 toolchains accidentally selecting unsupported native architectures on newer cards.
Runtime Layers¶
The fused path has three layers:
| Layer | Files | Role |
|---|---|---|
| Zero-copy views | include/diff/tensor_views.h, src/diff/tensor_views.cpp |
Wrap selected Model and SimState DeviceArray buffers as nova::Tensor external views. This keeps layout conversion out of physics kernels. |
| Differentiable collision forces | include/collision/diff_soft_contacts.h, src/collision/diff_soft_contacts.cpp |
Generated indexed Nova kernels that consume regular Contacts.soft_contact_* buffers. Current differentiable coverage is particle contact against world-owned static shapes (shape_body < 0); body-attached static or kinematic shapes are outside this operator's contract. |
| Differentiable particle dynamics | include/dynamics/semi_implicit/diff_particles.h, src/dynamics/semi_implicit/diff_particles.cpp |
Generated indexed Nova kernels for particle semi-implicit integration and point loss. |
| Differentiable rigid dynamics | include/dynamics/semi_implicit/diff_rigid.h, src/dynamics/semi_implicit/diff_rigid.cpp |
Generated indexed Nova kernels for static box-plane contact accumulation, rigid-body semi-implicit integration, and rigid terminal losses. |
| Shared traced contact formula | include/collision/diff_soft_contact_trace.h |
Single source of truth for static particle-shape contact math used by force-only and fused-step kernels. |
Non-differentiable collision broad phase, topology, and contact bookkeeping
should continue to use src/collision/collision_pipeline.cpp. Differentiable
force accumulation consumes the resulting regular Contacts and the same
Model/SimState storage through tensor views rather than copying into a
second simulation model.
Python API¶
Import Nova runtime and fused differentiable helpers with:
import novaphy
from novaphy import nova
import novaphy.diff as ndiff
from novaphy.solvers import DiffSemiImplicitSolver, DiffSolverConfig
novaphy.solvers is the canonical solver namespace. Generated operations and
tensor-view helpers remain under novaphy.diff; differentiable solver types
are intentionally not re-exported from the top-level novaphy namespace.
Tensor Views¶
| API | Purpose |
|---|---|
ndiff.state_particle_q_tensor(state, requires_grad=False) |
Zero-copy view of SimState.particle_q. |
ndiff.state_particle_qd_tensor(state, requires_grad=False) |
Zero-copy view of SimState.particle_qd. |
ndiff.state_particle_f_tensor(state, requires_grad=False) |
Zero-copy view of SimState.particle_f. |
ndiff.state_body_q_tensor(state, requires_grad=False) |
Zero-copy {body_count, 7} view of SimState.body_q. |
ndiff.state_body_qd_tensor(state, requires_grad=False) |
Zero-copy {body_count, 6} view of SimState.body_qd. |
ndiff.state_body_f_tensor(state, requires_grad=False) |
Zero-copy {body_count, 6} view of SimState.body_f. |
ndiff.state_joint_q_tensor(state, requires_grad=False) |
Zero-copy view of SimState.joint_q. |
ndiff.state_joint_qd_tensor(state, requires_grad=False) |
Zero-copy view of SimState.joint_qd. |
ndiff.model_tet_activation_tensor(model, requires_grad=False) |
Zero-copy view of Model.particle_tet_activation. |
ndiff.model_tet_material_tensor(model, requires_grad=False) |
Zero-copy view of Model.particle_tet_material. |
ndiff.model_tri_activation_tensor(model, requires_grad=False) |
Zero-copy view of Model.particle_tri_activation. |
ndiff.model_tri_material_tensor(model, requires_grad=False) |
Zero-copy view of Model.particle_tri_material. |
These views share storage with the original SimState. Set
requires_grad=True only for data that should receive adjoints.
Generated Physics Kernels¶
| API | Purpose |
|---|---|
ndiff.eval_static_particle_shape_contact_forces_generated(model, particle_q, particle_qd, particle_f, soft_contact_margin=0.0) |
Returns particle forces after world-owned static box/plane soft contacts are accumulated into particle_f. |
ndiff.assign_tet_materials_generated(material_params, base_tet_materials, tet_count) |
Builds per-tet material tensors from either 2 shared material parameters or 2 parameters per tet, preserving damping from the base material tensor. |
ndiff.eval_tetrahedra_forces_generated(model, particle_q, particle_qd, particle_f, tet_activations, tet_materials) |
Returns particle forces after Newton-style tetrahedral FEM forces are accumulated into particle_f. |
ndiff.eval_triangle_forces_generated(model, particle_q, particle_qd, particle_f, tri_activations, tri_materials) |
Returns particle forces after Newton-style triangle membrane/area/drag/lift forces are accumulated into particle_f. |
ndiff.integrate_particles_generated(model, particle_q, particle_qd, particle_f, dt) |
Returns (particle_q_next, particle_qd_next) using semi-implicit particle integration. |
ndiff.integrate_particles_with_static_contacts_generated(model, particle_q, particle_qd, particle_f, soft_contact_margin, dt) |
Returns (particle_q_next, particle_qd_next) after accumulating world-owned static box/plane soft contacts and integrating particles inside one generated indexed kernel. |
ndiff.integrate_rigid_bodies_with_static_contacts_generated(model, body_q, body_qd, body_f, angular_damping, friction_smoothing, dt) |
Returns (body_q_next, body_qd_next) after static box-plane penalty contacts and rigid semi-implicit integration in one generated indexed kernel. |
ndiff.body_twist_from_omega_generated(omega, linear_x, linear_y, linear_z) |
Builds packed rigid twists from differentiable angular velocity and fixed linear velocity. |
ndiff.rigid_orientation_velocity_loss_generated(...) |
Generic two-axis orientation and optional height loss with terminal linear/angular velocity regularization. |
ndiff.particle_point_loss_generated(particle_q, particle_index, target_x, target_y, target_z) |
Returns a scalar squared-distance loss for one particle. |
ndiff.particle_centroid_generated(particle_q) |
Returns the unweighted {1, 3} centroid of the particle positions. |
ndiff.vec3_point_loss_generated(point, target_x, target_y, target_z) |
Returns a scalar squared-distance loss from a vec3 tensor to a target. |
ndiff.clamp_f32_generated(x, lower, upper) |
Returns a graph-capturable clamped tensor for projected parameter updates. |
ndiff.record_planar_2r_featherstone_step(model, state_in, control, q, qd, q_out, qd_out, dt, angular_damping) |
Records the guarded analytic VJP for one supported, contact-free planar 2R SolverFeatherstone step. |
ndiff.planar_2r_tip_loss(...) |
Analytic end-effector squared-distance loss for the planar 2R articulation. |
ndiff.integrate_push_box_generated(...) |
Differentiable one-axis sphere-pusher/box penalty contact and box translation step. |
The kernels use nova::make_indexed_expr_kernel_typed_access_n and
nova::make_indexed_fixed_reduce_expr_kernel_typed_access_n, so model arrays
such as radii, flags, transforms, material parameters, and shape metadata are
read as source tensors. Dynamic state tensors such as particle_q and
particle_qd can require gradients.
Soft-body rest triangles and tetrahedra must be nondegenerate; ModelBuilder
rejects near-zero rest area/volume and inverted tetrahedron orientation.
Generated FEM treats zero volumetric stiffness as its finite zero-force limit
instead of evaluating a 0 / 0 material ratio.
Generated particle/contact/FEM vector rows are tightly packed {n, 3};
material rows likewise use their documented exact widths. Wider matrices are
rejected because these generated expressions use fixed scalarized row widths,
not arbitrary tensor strides.
Capture Graph Pattern¶
nova.ScopedCapture() records Nova tensor/kernel launches whose storage and
launch sequence stay fixed. It does not record the DeviceArray clear, copy,
snapshot, host staging, or component-mirror operations performed by a complete
DiffSemiImplicitSolver.step(). Consequently that solver reports
supports_graph_capture=False and must not be wrapped in Nova logical capture.
Pure Nova operator sequences, including tape.backward(), tensor gradient
clearing, and optimizer kernels, may still use ScopedCapture when all their
inputs and outputs are fixed Nova tensors. Driver-layer capture remains the
authoritative mechanism for engine operations that explicitly support it.
For vec3 parameters that need the same finite-gradient and norm clipping used by Newton's dice example:
For performance, enable Nova JIT before constructing generated kernels:
The diff demos do this by default. If Nova was built without
NOVA_ENABLE_NVRTC, the same code falls back to the bytecode interpreter.
Diffsim Ball Demo¶
python/demos/diff/demo_diffsim_ball.py follows Newton's
example_diffsim_ball.py framework path:
- The scene is built through the regular
ModelBuilder. - The ball is a regular particle in a regular
Model/SimState. - The floor is a regular static Z-up plane shape and the vertical wall is a regular static box shape.
CollisionPipeline.collide(states[0], contacts)runs once outside the Tape, matching Newton's one-shot static wall/ground contact generation.- Every substep calls
states[t].clear_forces()andDiffSemiImplicitSolver.step(states[t], states[t + 1], control, contacts, dt). - The solver consumes
Contacts.soft_contact_*, evaluates contact force in one generated kernel, and integrates particles in a second generated kernel. - Backward uses the generic Nova tape and generated adjoints, not a demo-specific backward path.
- The GUI trajectory is refreshed every few training iterations, so the drawn path follows the optimized ball trajectory.
Run a headless smoke test:
Expected behavior is a decreasing loss and changing optimized initial velocity.
The initial 289-substep trajectory and final velocity compare bit-for-bit with
Newton on the current CUDA development build. AD versus central finite
difference has a maximum relative error of about 2.2e-3.
Use --dump-states path.npz with --headless 0 to write the initial rollout
for Newton/NovaPhy comparison. Passing --headless N first performs N
training iterations, then dumps the optimized trajectory. The NPZ contains
frame-level particle_q shaped {sim_steps + 1, particle_count, 3} and also
particle_q_substeps shaped
{sim_steps * sim_substeps + 1, particle_count, 3}.
Diffsim Soft Body Demo¶
python/demos/diff/demo_diffsim_soft_body.py is a generated soft-body
material-optimization surrogate matching Newton's
example_diffsim_soft_body.py workflow:
- The soft grid, Z-up ground plane, and wall are built through the regular
ModelBuilder. - The optimized material parameters are stored as a Nova tensor and expanded to
per-tet material tensors by
assign_tet_materials_generated. - Each substep accumulates triangle forces, tetrahedral FEM forces, static particle-shape soft contact forces, and then integrates particles.
- The loss is the squared distance between particle center of mass and the target.
- Backward uses the generic Nova tape and generated adjoints.
The model and state storage are shared with the regular engine, but this rollout does not call production VBD. Its gradient therefore belongs to the generated FEM/contact/integration sequence listed above.
Run a headless smoke test:
Use --dump-states path.npz --headless 0 to write the initial rollout for
Newton/NovaPhy comparison. Passing --headless N first performs N training
iterations, then dumps the optimized trajectory.
Diffsim Dice Demo¶
python/demos/diff/demo_diffsim_dice.py is a smooth generated surrogate of
Newton's example_diffsim_dice.py rigid optimization workflow:
- The dice and Z-up ground are regular
ModelBuildershapes. - Candidate shape pairs come from the regular model's
shape_contact_pairs, the same explicit broad-phase topology consumed byCollisionPipeline. - Newton's mass construction is preserved: the body starts with the requested 4 kg base mass and the box contributes its default 1000 kg/m3 density.
- Static plane-box narrow phase keeps the four deepest box corners, applies shape margins to the contact points, and uses the same averaged penalty material, Huber-smoothed friction, gyroscopic term, and semi-implicit rigid integration.
- The generated rigid integrator uses the full body-frame inverse inertia, including world/body coordinate transforms and the gyroscopic term.
- One generated kernel performs all contact force/torque accumulation and integration for one substep. Nova generates its adjoint; there is no dice-specific backward implementation.
- Forward/backward kernels run as ordinary Nova launches. The host-side line search evaluates candidate trajectories with the regular forward rollout; this demo does not place the full training step in one CUDA graph.
- ViewerGL replays each refreshed trajectory frame by frame and retains prior training trajectories.
This demo reuses production model data and broad-phase topology, but it does
not differentiate a CollisionPipeline plus SolverSemiImplicit call. Its
gradient is the derivative of the generated box-plane penalty/contact rollout.
Run:
conda run -n novaphy_diff python python/demos/diff/demo_diffsim_dice.py
conda run -n novaphy_diff python python/demos/diff/demo_diffsim_dice.py --headless 10
conda run -n novaphy_diff python python/demos/diff/demo_diffsim_dice.py --check-grad
The contact-gradient check uses a deliberately penetrating box-plane state away from the discrete contact-generation threshold.
Diffsim Featherstone Demo¶
python/demos/diff/demo_diffsim_featherstone.py ports Newton's
example_diffsim_feathstone.py initial-hinge-velocity optimization:
- the forward rollout is the regular CUDA
SolverFeatherstoneoperating on ordinaryModelandSimStateobjects; joint_qandjoint_qdenter Nova through cached zero-copy tensor views;- every contact-free step records a hand-derived analytic VJP for the two-link planar revolute chain (mass matrix, Coriolis and gravity terms, joint limits, semi-implicit integration, and angular damping);
- the viewer keeps the arm in its vertical XOZ plane, colors the links blue and orange, renders a black-and-white checkerboard horizontal ground, and draws the end-effector trajectory in red;
- no
collision_novaorsim_novacompatibility object is used.
Run the optimization or the full 576-substep gradient check with:
conda run -n novaphy_diff python python/demos/diff/demo_diffsim_featherstone.py --headless 20
conda run -n novaphy_diff python python/demos/diff/demo_diffsim_featherstone.py --check-grad --headless 0
The analytic recorder rejects unsupported topology and solver features
explicitly. Its supported domain is one contact-free, world-rooted two-link
articulation with two enabled +Y revolute joints, dynamic bodies, identity
joint rotations, link COMs at their body origins, zero target drives,
friction, joint damping, feedforward force, and external body force. It also
checks q_out / qd_out against the analytic update before recording the VJP,
so an active effort/velocity clamp, contact, or other unmodeled forward effect
fails instead of silently returning a surrogate gradient.
Differentiable Ant Policy Fine-Tuning¶
python/demos/diff/demo_diffsim_ant_walk.py combines the D.VA Ant morphology,
a pretrained SB3 SAC Ant-v3 actor, the regular Featherstone/PGS forward solver,
and an analytic direct-force sensitivity. The default playback remains a pure
forward-policy demo. Passing --train N fine-tunes the eight biases in the
actor's final pre-tanh layer before replay.
The training gradient crosses every 1 ms Featherstone substep in the selected multi-frame horizon. It uses generalized mass matrices, spatial Jacobians, and a projection into the active contact-normal null space. Contact creation, state-dependent mass/bias derivatives, friction tangents, and gradients across the 20 Hz observation refresh are frozen. A real forward line search guards this approximation: a candidate is retained only if the ordinary solver measures a lower trajectory loss.
conda run -n novaphy_diff python \
python/demos/diff/demo_diffsim_ant_walk.py --check-gradient
conda run -n novaphy_diff python \
python/demos/diff/demo_diffsim_ant_walk.py \
--train 10 --train-horizon 8 \
--save-policy /tmp/novaphy_ant_trained.npz --headless 200
conda run -n novaphy_diff python \
python/demos/diff/demo_diffsim_ant_walk.py \
--policy /tmp/novaphy_ant_trained.npz
From-zero periodic gait¶
python/demos/diff/demo_diffsim_ant_from_scratch.py removes the pretrained
SAC policy. Its
80-parameter Actor starts at exact zeros, so the initial rollout applies no
joint torque. A horizon curriculum first optimizes standing and 4, 8, 12, 16,
24, 36, and 60 policy-frame trajectories, followed by a 400-frame
stabilization stage. The Actor combines first/second-harmonic phase features
with torso height, velocity, and tilt feedback. Its objective additionally
penalizes torso tilt/heading error, lateral drift, joint-pose distortion,
action discontinuity, and violations of diagonal-leg symmetry.
# Train from zero and open ViewerGL when training finishes.
conda run -n novaphy_diff python \
python/demos/diff/demo_diffsim_ant_from_scratch.py
# Replay the checked-in result without training.
conda run -n novaphy_diff python \
python/demos/diff/demo_diffsim_ant_from_scratch.py \
--load python/demos/data/ant/rl_policies/periodic_ant_natural_gait.npz
The saved natural-gait run remains upright for a 600-frame/30-second validation, travels 0.792 m, ends with positive forward velocity, and limits maximum lateral drift to 0.253 m and torso tilt to 5.32 degrees. This is still a slow nominal-terrain gait rather than a robust locomotion policy: no disturbance or terrain randomization is included.
Hybrid Franka Push-Box Demo¶
python/demos/diff/demo_diffsim_franka_push_box.py demonstrates a hybrid
gradient boundary on the repository's FR3 asset:
- regular CPU
SolverFeatherstoneadvances the seven-axis arm; novaphy.solvers.FeatherstoneAnalyticControlSensitivitypropagates an arbitrary controller parameterization through the generalized mass matrix, implicit PD response, and spatial point Jacobian;- the end-effector position/velocity trajectory is passed to
integrate_push_box_generated, whose smooth penalty contact and one-axis box integration are differentiated by Nova Tape; - the contact trajectory adjoints are multiplied by the analytic Featherstone trajectory Jacobian to obtain the key-pose gradient.
The reusable control-sensitivity backend supports a selected articulation
containing fixed and one-DOF joints. It is a frozen-dynamics linearization:
collision active-set changes and derivatives of state-dependent mass/bias
terms are not yet a complete general SolverFeatherstone.step_vjp. The demo
also assumes a stiff position-controlled pusher, so the generated contact
operator moves the box but does not feed its reaction force back into the arm.
Run the GUI, headless training, or directional gradient check with:
conda run -n novaphy_diff python \
python/demos/diff/demo_diffsim_franka_push_box.py
conda run -n novaphy_diff python \
python/demos/diff/demo_diffsim_franka_push_box.py --train 5
conda run -n novaphy_diff python \
python/demos/diff/demo_diffsim_franka_push_box.py --check-gradient
Current Coverage¶
| Area | Status |
|---|---|
Zero-copy SimState particle position/velocity/force views |
Implemented. |
Zero-copy SimState rigid pose/twist/force views |
Implemented. |
Zero-copy Model particle and static-shape source views |
Implemented internally. |
| World-owned static particle-vs-box soft contact | Differentiable generated path implemented. |
| World-owned static particle-vs-plane soft contact | Differentiable generated path implemented. |
| One-axis sphere-pusher/box contact | Differentiable generated path implemented for the hybrid Franka demo. |
| Generic Featherstone PD control sensitivity | Analytic frozen-dynamics backend implemented for selected fixed/one-DOF articulations. |
| Fused particle static-contact integration | Differentiable generated path implemented with fixed source-domain reduction and NVRTC JIT. |
| Tetrahedral FEM force | Differentiable generated path implemented for the Newton eval_tetrahedra formula. |
| Triangle membrane/area/drag/lift force | Differentiable generated path implemented for the Newton eval_triangle formula. |
| Soft-body material assignment | Differentiable generated path implemented for shared or per-tet mu/lambda parameters. |
| Particle semi-implicit integration | Differentiable generated path implemented. |
| Point target loss | Differentiable generated path implemented. |
| Particle center-of-mass loss building blocks | Differentiable generated path implemented. |
| Dynamic box-vs-static-plane rigid contact | Differentiable generated path implemented, including four-contact reduction and semi-implicit integration. |
| General dynamic rigid body contacts | Not fused yet. Use the shared non-differentiable collision pipeline for unsupported topology/bookkeeping. |
| Mesh contacts | Not fused yet. |
| Spring and bending forces | Not fused yet. |
| Soft-body static primitive shape contacts | Differentiable generated path implemented for particle-vs-box/plane. |
| Soft-body dynamic/mesh contact topology | Not fused yet; collision topology still needs to come from the shared collision pipeline. |
| VBD | Optional module; can be fully disabled with NOVAPHY_BUILD_VBD=OFF. |
Design Rules¶
- Keep physics logic shared with the forward engine wherever possible.
- Keep zero-copy adaptation in
diff/tensor_views.*. - Put differentiable math in domain files beside the forward code, such as
collision/diff_soft_contacts.*anddynamics/semi_implicit/diff_particles.*. - Prefer generated indexed Nova kernels for differentiable forward and adjoint code.
- Use
collision_pipeline.cppfor non-differentiable broad phase and contact topology instead of creating a second collision stack. - Do not add new dependencies on
collision_novaorsim_nova.