LATX: optimize scalar pixel updates and VROUNDPS truncation - #442
LATX: optimize scalar pixel updates and VROUNDPS truncation#442luzeng87 wants to merge 12 commits into
Conversation
Track VEX.128 destinations whose architectural YMM high halves are known to be zero, and materialize those clears only when a 256-bit operation can observe them or before leaving the TB. This removes repeated LASX clear instructions while preserving signal, JIT, TU, and AOT-visible state. Add a standalone JIT, cold-AOT, and hot-AOT semantic test for the deferred state. Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Extend high-bit reduction to VEX scalar arithmetic, FMA, and move instructions, and emit scalar results directly when the preserved lanes are dead. Remove redundant vector temporaries and copies from compare, bitwise, shuffle, min/max, and packed FMA translations. Use LSX bit selection for packed min/max results. On Geekbench 7 Audio Encoder, these changes are part of the hot-AOT improvement from 557 to 596 before scalar FMA forwarding. Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Recognize scalar add, subtract, or multiply results that are immediately inserted into an XMM destination and then consumed by scalar FMADD or FMSUB. Keep the scalar value in its temporary until the fused operation and remove the intermediate VEXTRINS.W. Require exact operand and insertion matches so unrelated IR2 sequences remain unchanged. Add a standalone test for FMADD, FMSUB, NaN payloads, and preserved XMM upper lanes. The optimization raised Geekbench 7 Audio Encoder hot AOT from 596 to a 600 median. Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
|
This stacked branch includes the unsafe deferred YMM-state change from #437. Keeping its unrelated instruction changes in one branch also makes correctness attribution difficult. Closing it; only independently verified functional changes should remain open. |
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
|
Synced the YMM mask/blend writeback correction as commit |
|
Superseded by #438, which contains the scalar pixel update and VROUNDPS changes together with the required alias fix and instruction-based tests. |
Summary
This PR is stacked on #437, #438, and #439; the scalar pixel-update and VROUNDPS changes are the final commits. The CPU feature-reporting change remains separate in #440.
Validation
Signed-off-by: Lu Zeng luzeng87@gmail.com