Source audit and StateFlags fixes — 2026-09-08¶
This change addresses three historical source checks and one upstream Python API
compatibility failure on top of PR preparation commit 3014386. It does not
change MPM formulas, CUDA kernels, numerical tolerances or upstream tests.
The validation below was performed in the independent PR checkout before
publication. The user subsequently authorized committing and pushing these
fixes to the personal pr/four-scene11-20260908 branch only; no upstream PR or
default-runtime switch is part of that publication.
Findings and changes¶
Legacy Jp source anchor¶
The old inline clampf(p.Jp * oldJ / det_new, ...) became a call to
update_jp_exp_tr_plastic in the shared material header. This helper still uses
the old determinant ratio when hardening and softening rates are both one, and
clamps the resulting Jp. The newer official plastic-trace route is separate.
The D1 audit now checks the live helper call, the legacy-route guard, the ratio,
the unit-rate condition and the result clamp. Four tests exercise the compiled
helper for compression, expansion and both clamp bounds against explicit
expected values. A copied-source negative probe removing the helper call fails
the D1 audit at NP-g2p-jp. Historical D1 comparison documents remain historical;
this audit refresh does not certify all legacy formulas as Newton parity.
Mass/momentum scaling source inventory¶
The old standalone scale_grid_mass_momentum_inv_v launch is no longer the
active route. Finite mass is scaled inside P2G; momentum uses that scaled mass.
CPU dense/world paths and CUDA dense/world host paths supply the inverse volume,
and kinematic sentinel mass retains its separate handling.
The source test follows these live CPU/host/kernel sites and keeps the check
that classic route C does not opt into the NewtonFEM flag. Negative cases replace
CUDA mass scaling with multiplication by one and must fail the audit. Existing
numerical scaling and kinematic-sentinel tests are retained. The separate
test_step5_delassus_flow_magnitude_gates failure is outside this repair and
remains in the full suite.
GS shared-memory inventory¶
The only __shared__ declarations in rheology_colored_gs.cu belong to
k_tiled_rho_sum_max, the separate residual sum/max reduction. They are not in
the per-cell GS solve kernels. The old whole-file prohibition predated that
reduction.
The test now bounds the exception to this named function, checks its sum/max
scratch and synchronization, and rejects shared memory anywhere else in the
file. The no-SVD, kernel-presence, launch-geometry and residual checks remain.
A negative test injecting shared memory into k_nf_gs_color must fail.
StateFlags compatibility and migration¶
The Python interface now exposes exactly the seven members required by upstream
753894f: JOINT_Q, JOINT_QD, BODY_Q, BODY_QD, PARTICLE_Q, PARTICLE_QD,
and ALL. Values are unchanged. Only Python PARTICLE / Particle aliases
are removed. The C++ aggregate bit mask and reset implementation are unchanged.
No product/demo caller of these Python aliases was found in the PR checkout.
External callers using a snapshot alias should migrate explicitly:
particle_flags = int(novaphy.StateFlags.PARTICLE_Q) | int(novaphy.StateFlags.PARTICLE_QD)
solver.reset(state, flags=particle_flags)
For this MPM reset implementation, either particle bit selects rest-history reset; it does not overwrite position or velocity. Six new cases cover default, combined, individual particle, body-only and zero flags, checking all five MPM history arrays and unchanged particle kinematics. The upstream exact-membership assertion is unchanged.
Verification¶
The four original failures were reproduced before edits. Both native binding files were backed up; edits are limited to enum exports and the reset docstring. The rebuilt standard wheel is installed in a separate venv, preserving the old PR wheel and environment. CUDA configuration is unchanged (MPM and DIFF ON, IPC and Featherstone CUDA OFF, sm86).
- Wheel SHA256:
ee9ad97781a2177c04e2cbe26ac8606f85bfa2305168e351eb48ee632a267c7a. - All 217 packaged Python files match the current source.
- Focused source/Jp/scaling/API/reset group: 34 passed, one known unrelated step5 numerical failure deselected from this group and retained in full-suite execution.
- Negative source checks cover removed Jp updates, missing CUDA mass scaling and forbidden GS shared memory.
- Full suite: 1883 passed, 16 failed, 132 skipped, 0 errors, 656.88 seconds. The four repaired targets pass in this full run.
- MPM lifecycle/binding/reset memcheck: 30 passed, zero reported errors.
- Four Viser servers: each HTTP 200, ten frame updates and normal exit. Four two-frame sampled-state sequences: four passed under memcheck, zero reported errors.
- Changed C++ lines pass clang-format dry run; no new Ruff findings. clang-tidy is unavailable, so no clang-tidy certification is claimed.
- Remaining failed nodes are listed in the machine-readable summary. No new failed node compared with the original full-suite run: True. This is not whole-engine acceptance while failures remain.
Tests renamed to reflect their actual scope:
| Previous test | Replacement |
|---|---|
test_step4_scale_launch_sites_are_newton_fem_only |
test_step4_mass_scale_is_fused_into_newton_fem_p2g_only |
test_novaphy_gs_cu_no_shared_no_svd |
test_novaphy_gs_shared_only_in_residual_reduction_no_svd |
The D1 test and the upstream StateFlags test keep their original names.