Consistent by construction
Every projection sharing an activation site uses the same transform. Calibration, exported weights, and runtime stay in the same coordinates.
Native Low-Bit Quantization of Vision-Language-Action Models via Consistent Folding
1.20–1.33×
Native W4A4 vs the TRT BF16 engine, end to end.
GR00T N1.5 / N1.6 / N1.7 and π0.5 · batch 1
2.11×
5,325 → 2,525 MB of serialized TensorRT plans.
GR00T N1.6 · not peak GPU memory.
92.5%
W4A4 + o/d INT8 on Jetson, GR00T N1.7, four tasks.
BF16 PyTorch 87.5% · TRT BF16 90.0% · W4A4 80.0%.
0.460.85
W4A4 → W4A4 + o/d INT8, GR00T N1.6 held-out.
Costs +0.4–1.8 ms end to end on GR00T desktop.
Overview
What does lower precision actually buy a robot policy?
FoldQuantVLA fixes a shared change of basis at each activation site, folds it into weights and normalization gains, and rounds once. A common build path produces floating-point, eight-bit, and four-bit engines for the language backbone and iterative action expert.
We evaluate the resulting engines with compiled floating-point controls and 48 closed-loop LIBERO campaigns. The controls separate deployment speedups from precision gains; the rollouts show where offline action fidelity stops predicting task success.
Every projection sharing an activation site uses the same transform. Calibration, exported weights, and runtime stay in the same coordinates.
Custom TensorRT plugins fuse activation transforms and quantization with integer tensor-core GEMMs on desktop and Jetson GPUs.
Compiled controls, device measurements, and closed-loop outcomes reveal both the gains and the configurations that fail.
Method
One transform per activation site. One engine for the denoising loop.
Scaling and rotation define the coordinates once. Every consumer of the activation must agree on that basis before rounding.
Language and expert projections use W8A8 or W4A4. Vision, attention softmax, residuals, and decoding remain floating point.
Dynamic per-token scaling serves each denoising step. Calibration optimality across steps requires the paper’s stated directional-distribution condition.
Closed-loop simulation · LIBERO & SimplerEnv
FoldQuantVLA against low-bit baselines in two simulators, LIBERO and SimplerEnv, with the BF16 reference first.
The note under each table gives the protocol behind every row.
LIBERO four suites · 200 episodes per suite
| Arm | Prec. | Spatial | Object | Goal | Long | Average SR ↑ |
|---|---|---|---|---|---|---|
| BF16 PyTorch | bf16 | 98.5% | 98.5% | 92.5% | 93.5% | 95.75% |
| HoloQ-VLA† | int4 | 94.0% | 98.5% | 89.5% | 90.0% | 93.00% |
| DuQuant† | int4 | 97.0% | 97.5% | 95.0% | 83.5% | 93.25% |
| FoldQuant W4A4 (ours) | int4 | 98.5% | 97.0% | 95.5% | 88.5% | 94.88% |
| FoldQuant W4A4 + o/d INT8 (ours) | int4 | 96.5% | 98.5% | 94.0% | 93.5% | 95.62% |
FoldQuant rows are real INT4 / INT8 TensorRT engines. † HoloQ-VLA and DuQuant are our emulated implementations of those methods in the same release repository.
LIBERO four suites · 200 episodes per suite
| Arm | Prec. | Spatial | Object | Goal | Long | Average SR ↑ |
|---|---|---|---|---|---|---|
| BF16 PyTorch | bf16 | 98.5% | 99.0% | 97.5% | 93.5% | 97.12% |
| HoloQ-VLA | int4 | 99.0% | 97.0% | 100.0% | 96.0% | 98.00% |
| DuQuant | int4 | 96.0% | 99.0% | 94.0% | 88.0% | 94.25% |
| QuantVLA | int4 | 94.0% | 98.0% | 80.0% | 56.0% | 82.00% |
| SmoothQuant | int4 | 83.0% | 88.0% | 40.0% | 26.0% | 59.25% |
| FoldQuant W4A4 (ours) | int4 | 98.0% | 99.5% | 96.0% | 95.5% | 97.25% |
| FoldQuant W4A4 + o/d INT8 (ours) | int4 | 99.5% | 99.0% | 98.5% | 94.0% | 97.75% |
FoldQuant rows are real INT4 / INT8 TensorRT engines. BF16 and FoldQuant rows: FoldQuant released openpi harness, 20 trials per task, K = 5, seed 7. HoloQ-VLA, DuQuant, QuantVLA and SmoothQuant: reported by HoloQ-VLA (emulated quantization, 10 calibration samples); their FP16 reference is also 97.1%. Our reproduction of HoloQ-VLA W4A4 on the long suite reaches 94.0% against 96.0% reported (t = −0.82, p = 0.44).
LIBERO · four suites · 40 tasks · 800 episodes per configuration
| Checkpoint | H | K | BF16 SR | FoldQuant W8A8 | FoldQuant W4A4 | FoldQuant W4A4 + o/d INT8 | ||||
|---|---|---|---|---|---|---|---|---|---|---|
| PyTorch | TensorRT | SR ↑ | Median cos ↑ | SR ↑ | Median cos ↑ | SR ↑ | Median cos ↑ | |||
| GR00T N1.7 | 16 | 8 | 96.25% | 95.50% | 95.38% | 0.99996 | 95.38% | 0.99817 | 95.00% | 0.99928 |
| GR00T N1.6 | 16 | 8 | 96.38% | 97.75% | 95.75% | 0.99998 | 95.75% | 0.99876 | 96.62% | 0.99967 |
| GR00T N1.5 | 16 | 1 | 86.38% | 86.00% | 87.12% | 0.99999 | 87.38% | 0.99885 | 87.00% | 0.99933 |
| π0.5 | 10 | 5 | 96.50% | 98.00% | 97.38% | 1.00000 | 97.12% | 0.99942 | 97.62% | 0.99981 |
H is the action chunk length and K the executed prefix.
Seven WidowX tasks · 200 episodes per task
| Arm | Task | Avg ↑ | ||||||
|---|---|---|---|---|---|---|---|---|
| Put spoon on towel | Put carrot on plate | Put eggplant in basket | Stack green cube on yellow | Put eggplant in sink | Close drawer | Open drawer | ||
| BF16 PyTorch | 63.5% | 60.5% | 92.5% | 6.0% | 43.0% | 72.0% | 98.0% | 62.2% |
| FoldQuant W8A8 (ours) | 63.5% | 59.5% | 89.5% | 2.5% | 40.0% | 64.5% | 96.5% | 59.4% |
| FoldQuant W4A4 (ours) | 66.0% | 70.5% | 58.0% | 7.5% | 63.0% | 82.0% | 81.5% | 61.2% |
| FoldQuant W4A4 + o/d INT8 (ours) | 75.5% | 66.5% | 70.5% | 5.5% | 50.0% | 89.5% | 92.0% | 64.2% |
Real-world tasks
GR00T N1.7 deployed on Jetson AGX Orin: one single-arm ALOHA task and three SO-101 tasks, 20 episodes per task.
GR00T N1.7 · pick a task; every engine plays at once so their pace can be compared
Prompt“Pick blue cube and place on red cube”
One selected episode per engine in each scene · external camera · 4× speed
* ModelOpt W8A8 SmoothQuant and ModelOpt W4A16 AWQ are run with NVIDIA’s TensorRT Model Optimizer (ModelOpt) repository.
Prompt“Pick up the banana and place it in the pot, then close the lid”
One selected episode per engine in each scene · external camera · 4× speed
Prompt“Pick all cubes and place into cup”
One selected episode per engine in each scene · external camera · 4× speed
Prompt“Use the right gripper to pick up the banana and place it into the pot. Then pick up the lid with the right gripper and place it on top of the pot to close it.”
One selected episode per engine in each scene · high camera · 4× speed
GR00T N1.7 on Jetson AGX Orin · SR % · 20 episodes per task and arm · Wilson 95% interval
| Arm | ALOHA · Banana | SO101 · Banana | SO101 · All blocks into cup | SO101 · Blue block on red | Average SR ↑ |
|---|---|---|---|---|---|
| BF16 PyTorch | 90% | 90% | 100% | 70% | 87.5%95% CI [78.5, 93.1] |
| TRT BF16 (float engine) | 95% | 90% | 100% | 75% | 90.0%95% CI [81.5, 94.8] |
| ModelOpt W8A8 SmoothQuant* | N/A | N/A | N/A | 30% | 30.0%95% CI [14.5, 51.9]T4 only |
| ModelOpt W4A16 AWQ* | N/A | N/A | N/A | 40% | 40.0%95% CI [21.9, 61.3]T4 only |
| FoldQuant W8A8 (ours) | 85% | 95% | 100% | 70% | 87.5%95% CI [78.5, 93.1] |
| FoldQuant W4A4 (ours) | 75% | 85% | 95% | 65% | 80.0%95% CI [70.0, 87.3] |
| FoldQuant W4A4 + o/d INT8 (ours) | 95% | 95% | 100% | 80% | 92.5%95% CI [84.6, 96.5] |
Arms interleaved with matched object placements. W4A4 + o/d INT8 vs W4A4: +20 / +10 / +5 / +15 pp per task, +12.5 pp on average. ModelOpt W8A8 SmoothQuant stopped after one task (SO-101 blue on red): jerky motion risked the hardware. It is not in any total. ModelOpt W4A16 AWQ was also run on that task only: 8/20 (40%). Median logged-observation action cosine vs BF16 PyTorch — ALOHA: TRT BF16 0.99999 · W8A8 0.99999 · W4A4 0.99962 · W4A4 + o/d INT8 0.99975; SO-101: TRT BF16 1.00000 · ModelOpt SmoothQuant 0.99892 · W8A8 1.00000 · W4A4 0.99961 · W4A4 + o/d INT8 0.99988. * ModelOpt W8A8 SmoothQuant and ModelOpt W4A16 AWQ are run with NVIDIA’s TensorRT Model Optimizer (ModelOpt) repository.
π0.5 · every engine plays at once so their pace can be compared
Prompt“Pick blue cube and place on red cube”
One selected episode per engine in each scene · external camera · 4× speed
π0.5 on Jetson AGX Orin · SR % · 20 episodes per arm · Wilson 95% interval
| Arm | SR ↑ |
|---|---|
| BF16 PyTorch | 80.0%95% CI [58.4, 91.9] |
| TRT BF16 (float engine) | 100.0%95% CI [83.9, 100.0] |
| FoldQuant W4A4 (ours) | 85.0%95% CI [64.0, 94.8] |
20 episodes per arm on one task. Median logged-observation action cosine of W4A4 vs BF16 PyTorch is 0.98994; its backbone cosine (0.819) is the lowest measured on either robot while success stays at the reference. W8A8 and W4A4 + o/d INT8 were not run for π0.5 on the robot.
Results · Latency & memory
Pick a target, then a model family.
Every arm of that family in one table.
Four checkpoints · TensorRT 10.3 · batch 1 · observation-to-action chunk
| Arm | Prec. | GPU (ms) ↓ | E2E (ms) ↓ | Rate (Hz) ↑ | vs TRT BF16 ↑ |
|---|---|---|---|---|---|
| TRT BF16 (float engine) | bf16 | 119.1 | 146 | 6.8 | ref |
| Eager PyTorch | bf16 | 323.2 | 351 | 2.8 | 0.42× |
| torch.compile | bf16 | 202.4 | 231 | 4.3 | 0.63× |
| ModelOpt W8A8 SQ | int8 | 109.5 | 137 | 7.3 | 1.07× |
| ModelOpt W4A16 AWQ | int4 | 139.1 | 167 | 6.0 | 0.87× |
| FoldQuant W8A8 (ours) | int8 | 100.5 | 127 | 7.9 | 1.15× |
| FoldQuant W4A4 (ours) | int4 | 91.8 | 119 | 8.4 | 1.23× |
| FoldQuant W4A4 + o/d INT8 (ours) | int4 | 93.5 | 120 | 8.3 | 1.22× |
| Arm | Prec. | GPU (ms) ↓ | E2E (ms) ↓ | Rate (Hz) ↑ | vs TRT BF16 ↑ |
|---|---|---|---|---|---|
| TRT BF16 (float engine) | bf16 | 130.1 | 150 | 6.7 | ref |
| Eager PyTorch | bf16 | 313.2 | 333 | 3.0 | 0.45× |
| torch.compile | bf16 | 196.6 | 217 | 4.6 | 0.69× |
| ModelOpt W8A8 SQ | int8 | 121.5 | 142 | 7.0 | 1.06× |
| ModelOpt W4A16 AWQ | int4 | 155.6 | 177 | 5.6 | 0.85× |
| FoldQuant W8A8 (ours) | int8 | 117.0 | 137 | 7.3 | 1.09× |
| FoldQuant W4A4 (ours) | int4 | 104.8 | 125 | 8.0 | 1.20× |
| FoldQuant W4A4 + o/d INT8 (ours) | int4 | 106.3 | 127 | 7.9 | 1.18× |
| Arm | Prec. | GPU (ms) ↓ | E2E (ms) ↓ | Rate (Hz) ↑ | vs TRT BF16 ↑ |
|---|---|---|---|---|---|
| TRT BF16 (float engine) | bf16 | 107.0 | 135 | 7.4 | ref |
| Eager PyTorch | bf16 | 216.2 | 249 | 4.0 | 0.54× |
| torch.compile | bf16 | 166.1 | 199 | 5.0 | 0.68× |
| ModelOpt W8A8 SQ | int8 | 105.4 | 134 | 7.5 | 1.01× |
| ModelOpt W4A16 AWQ | int4 | 144.8 | 174 | 5.7 | 0.78× |
| FoldQuant W8A8 (ours) | int8 | 94.7 | 123 | 8.1 | 1.10× |
| FoldQuant W4A4 (ours) | int4 | 82.1 | 111 | 9.0 | 1.22× |
| FoldQuant W4A4 + o/d INT8 (ours) | int4 | 85.0 | 114 | 8.8 | 1.18× |
| Arm | Prec. | GPU (ms) ↓ | E2E (ms) ↓ | Rate (Hz) ↑ | vs TRT BF16 ↑ |
|---|---|---|---|---|---|
| TRT BF16 (float engine) | bf16 | 224.7 | 229 | 4.4 | ref |
| Eager PyTorch | bf16 | 856.9 | 861 | 1.2 | 0.27× |
| torch.compile | bf16 | 337.9 | 342.5 | 2.9 | 0.67× |
| ModelOpt W8A8 SQ | int8 | 206.1 | 211 | 4.7 | 1.09× |
| ModelOpt W4A16 AWQ | int4 | 321.6 | 326 | 3.1 | 0.70× |
| FoldQuant W8A8 (ours) | int8 | 221.1 | 226 | 4.4 | 1.01× |
| FoldQuant W4A4 (ours) | int4 | 167.6 | 172 | 5.8 | 1.33× |
| FoldQuant W4A4 + o/d INT8 (ours) | int4 | 198.7 | 203 | 4.9 | 1.13× |
Four checkpoints · TensorRT 10.15 · batch 1 · observation-to-action chunk
| Arm | Prec. | GPU (ms) ↓ | E2E (ms) ↓ | Rate (Hz) ↑ | vs TRT BF16 ↑ |
|---|---|---|---|---|---|
| TRT BF16 (float engine) | bf16 | 39.0 | 41.0 | 24.4 | ref |
| Eager PyTorch | bf16 | 65.0 | 67.3 | 14.9 | 0.61× |
| torch.compile | bf16 | N/A | 54.7 | 18.3 | 0.75× |
| ModelOpt W8A8 SQ | int8 | 34.0 | 37.5 | 26.7 | 1.09× |
| ModelOpt W4A16 AWQ | int4 | 52.2 | 60.3§ | 16.6 | 0.69× |
| FoldQuant W8A8 (ours) | int8 | 33.0 | 35.9 | 27.9 | 1.14× |
| FoldQuant W4A4 (ours) | int4 | 30.0 | 32.2 | 31.1 | 1.27× |
| FoldQuant W4A4 + o/d INT8 (ours) | int4 | 31.0 | 34.0 | 29.4 | 1.21× |
o/d INT8 vs W4A4: +1.8 ms E2E · § ModelOpt W4A16 AWQ measured in the framework runtime; its ratio uses that runtime’s TRT BF16 (41.5 ms) · GPU components printed as whole ms by the upstream timer. o/d INT8 uses the campaign recipe engine.
| Arm | Prec. | GPU (ms) ↓ | E2E (ms) ↓ | Rate (Hz) ↑ | vs TRT BF16 ↑ |
|---|---|---|---|---|---|
| TRT BF16 (float engine) | bf16 | 42.5 | 46.6 | 21.5 | ref |
| Eager PyTorch | bf16 | 64.7 | 68.8 | 14.5 | 0.68× |
| torch.compile | bf16 | N/A | 59.85§ | 16.7 | 0.74× |
| ModelOpt W8A8 SQ | int8 | 33.1 | 40.3§ | 24.8 | 1.09× |
| ModelOpt W4A16 AWQ | int4 | 57.2 | 64.1§ | 15.6 | 0.69× |
| FoldQuant W8A8 (ours) | int8 | 36.4 | 40.6 | 24.6 | 1.15× |
| FoldQuant W4A4 (ours) | int4 | 31.6 | 35.8 | 27.9 | 1.30× |
| FoldQuant W4A4 + o/d INT8 (ours) | int4 | 33.0 | 37.1 | 27.0 | 1.26× |
o/d INT8 vs W4A4: +1.3 ms E2E · § torch.compile, ModelOpt W8A8 SQ and ModelOpt W4A16 AWQ measured in the framework runtime; their ratio uses that runtime’s TRT BF16 (44.1 ms)
| Arm | Prec. | GPU (ms) ↓ | E2E (ms) ↓ | Rate (Hz) ↑ | vs TRT BF16 ↑ |
|---|---|---|---|---|---|
| TRT BF16 (float engine) | bf16 | 36.9 | 41.5 | 24.1 | ref |
| Eager PyTorch | bf16 | 50.6 | 54.8 | 18.2 | 0.76× |
| torch.compile | bf16 | N/A | 53.46§ | 18.7 | 0.75× |
| ModelOpt W8A8 SQ | int8 | 25.4 | 34.1§ | 29.3 | 1.18× |
| ModelOpt W4A16 AWQ | int4 | 43.9 | 53.2§ | 18.8 | 0.76× |
| FoldQuant W8A8 (ours) | int8 | 32.7 | 38.0 | 26.3 | 1.09× |
| FoldQuant W4A4 (ours) | int4 | 28.4 | 33.1 | 30.2 | 1.25× |
| FoldQuant W4A4 + o/d INT8 (ours) | int4 | 29.1 | 33.5 | 29.9 | 1.24× |
o/d INT8 vs W4A4: +0.4 ms E2E · § torch.compile, ModelOpt W8A8 SQ and ModelOpt W4A16 AWQ measured in the framework runtime; their ratio uses that runtime’s TRT BF16 (40.2 ms)
| Arm | Prec. | GPU (ms) ↓ | E2E (ms) ↓ | Rate (Hz) ↑ | vs TRT BF16 ↑ |
|---|---|---|---|---|---|
| TRT BF16 (float engine) | bf16 | 102.8 | 113.8 | 8.8 | ref |
| Eager PyTorch | bf16 | 154.9 | 165.7 | 6.0 | 0.69× |
| torch.compile | bf16 | N/A | 99.6 | 10.0 | 1.14× |
| ModelOpt W8A8 SQ | int8 | N/A | N/A | N/A | N/A |
| ModelOpt W4A16 AWQ | int4 | N/A | N/A | N/A | N/A |
| FoldQuant W8A8 (ours) | int8 | 77.7 | 88.2 | 11.3 | 1.29× |
| FoldQuant W4A4 (ours) | int4 | 64.5 | 74.7 | 13.4 | 1.52× |
| FoldQuant W4A4 + o/d INT8 (ours) | int4 | 70.3 | 80.9 | 12.4 | 1.41× |
o/d INT8 vs W4A4: +6.2 ms E2E · ModelOpt SQ / AWQ are not buildable on a 16 GB GPU
Four checkpoints · batch 1 · steady device memory during inference, one process per arm
| Arm | Prec. | Floor (MiB) ↓ | As served (MiB) ↓ | Weights (MB) ↓ | Floor vs TRT BF16 |
|---|---|---|---|---|---|
| TRT BF16 (float engine) | bf16 | 7,193 | 7,193 | 6,482 | ref |
| Eager PyTorch | bf16 | 6,469 | 6,469 | 6,288 | −10% |
| FoldQuant W8A8 (ours) | int8 | 6,149 | 6,149 | 4,782 | −15% |
| FoldQuant W4A4 (ours) | int4 | 5,563 | 5,563 | 3,689 | −23% |
| FoldQuant W4A4 + o/d INT8 (ours) | int4 | 5,513 | 5,513 | 3,866 | −23% |
The upstream N1.7 pipeline already deletes replaced modules, so floor equals as served.
| Arm | Prec. | Floor (MiB) ↓ | As served (MiB) ↓ | Weights (MB) ↓ | Floor vs TRT BF16 |
|---|---|---|---|---|---|
| TRT BF16 (float engine) | bf16 | 6,349 | 10,609 | 3,802 | ref |
| Eager PyTorch | bf16 | 6,947 | 6,947 | 6,574 | +9% |
| FoldQuant W8A8 (ours) | int8 | 5,259 | 9,519 | 2,104 | −17% |
| FoldQuant W4A4 (ours) | int4 | 4,673 | 8,933 | 1,011 | −26% |
| FoldQuant W4A4 + o/d INT8 (ours) | int4 | 4,623 | 8,883 | 1,189 | −27% |
| Arm | Prec. | Floor (MiB) ↓ | As served (MiB) ↓ | Weights (MB) ↓ | Floor vs TRT BF16 |
|---|---|---|---|---|---|
| TRT BF16 (float engine) | bf16 | 5,151 | 7,987 | 2,315 | ref |
| Eager PyTorch | bf16 | 5,797 | 5,797 | 5,448 | +13% |
| FoldQuant W8A8 (ours) | int8 | 4,451 | 7,287 | 1,282 | −14% |
| FoldQuant W4A4 (ours) | int4 | 4,111 | 6,947 | 635 | −20% |
| FoldQuant W4A4 + o/d INT8 (ours) | int4 | 4,059 | 6,895 | 736 | −21% |
| Arm | Prec. | Floor (MiB) ↓ | As served (MiB) ↓ | Weights (MB) ↓ | Floor vs TRT BF16 |
|---|---|---|---|---|---|
| TRT BF16 (float engine) | bf16 | 6,971 | 12,313 | 4,849 | ref |
| Eager PyTorch | bf16 | 7,691 | 7,691 | 7,473 | +10% |
| FoldQuant W8A8 (ours) | int8 | 4,927 | 10,269 | 2,443 | −29% |
| FoldQuant W4A4 (ours) | int4 | 4,553 | 9,895 | 1,349 | −35% |
| FoldQuant W4A4 + o/d INT8 (ours) | int4 | 4,651 | 9,993 | 1,669 | −33% |
What the controls reveal
The compiled float engine already supplies 61–75% of the eager-to-W4A4 reduction on the desktop and 83–92% on Jetson AGX Orin, where four-bit adds the last 8–17%. Reporting speedups against compiled float engines keeps graph optimisation from being credited to quantization.
When the weights an engine replaces are not kept on the GPU, FoldQuant W4A4 needs 20–35% less steady device memory than the TRT BF16 engine on all four checkpoints. Vision, embeddings and runtime workspaces set the floor that bit width cannot remove.
Above 0.99 held-out action cosine, success differences between arms stay within closed-loop noise and are not ordered by cosine. High cosine screens a broken configuration without ranking the working ones.
Citation
Anonymous manuscript · 2026
@misc{anonymous2026foldquantvla,
title = {FoldQuantVLA: Native Low-Bit Quantization of Vision-Language-Action Models via Consistent Folding},
author = {{Anonymous Authors}},
year = {2026},
note = {Anonymous ICRA submission}
}