Hello everyone,
We are currently experiencing several issues with the new VCU118 design provided in the 2026.2 release and were wondering if anyone has encountered something similar.
For context, we have a working design for the same FPGA based on the 2024.4 release. After migrating to the 2026.2 release, we started seeing the following problems:
-
L2-Lite enabled: We are unable to connect using GRMON (v4.1.2 64-bit evaluation version). GRMON reports:
ERROR! AMBA plug&play not found! -
L2 disabled: GRMON detects the CPU correctly, but execution immediately fails with a Load access fault. What makes this particularly confusing is that the exception occurs at PC = 0x00000002, even though compressed instructions are disabled and no load instruction is being executed at that point.
By single-stepping through the execution, we observed that execution unexpectedly jumps to the ROM address before the crash. The trace is shown below:
grmon4> step 50
0x0000000000000000: 0000b197 auipc gp, 0xb <_start+0>
0x0000000000000000: 0000b197 auipc gp, 0xb <_start+0>
0x00000000c0000000: 00000093 li ra, 0
0x00000000c0000004: 00000113 li sp, 0
0x00000000c0000008: 00000193 li gp, 0
0x00000000c000000c: 00000213 li tp, 0
0x00000000c0000010: 00000293 li t0, 0
0x00000000c0000014: 00000313 li t1, 0
0x00000000c0000018: 00000393 li t2, 0
0x00000000c000001c: 00000413 li s0, 0
0x00000000c0000020: 00000493 li s1, 0
0x00000000c0000024: 00000513 li a0, 0
0x00000000c0000028: 00000593 li a1, 0
0x00000000c000002c: 00000613 li a2, 0
0x00000000c0000030: 00000693 li a3, 0
0x00000000c0000034: 00000713 li a4, 0
0x00000000c0000038: 00000793 li a5, 0
0x00000000c000003c: 00000813 li a6, 0
0x00000000c0000040: 00000893 li a7, 0
0x00000000c0000044: 00000913 li s2, 0
0x00000000c0000048: 00000993 li s3, 0
0x00000000c000004c: 00000a13 li s4, 0
0x00000000c0000050: 00000a93 li s5, 0
0x00000000c0000054: 00000b13 li s6, 0
0x00000000c0000058: 00000b93 li s7, 0
0x00000000c000005c: 00000c13 li s8, 0
0x00000000c0000060: 00000c93 li s9, 0
0x00000000c0000064: 00000d13 li s10, 0
0x00000000c0000068: 00000d93 li s11, 0
0x00000000c000006c: 00000e13 li t3, 0
0x00000000c0000070: 00000e93 li t4, 0
0x00000000c0000074: 00000f13 li t5, 0
0x00000000c0000078: 00000f93 li t6, 0
0x00000000c000007c: 00010137 lui sp, 0x10
0x00000000c0000080: 00000297 auipc t0, 0x0
0x00000000c0000084: 02428293 addi t0, t0, 36
0x00000000c0000088: 30529073 csrw mtvec, t0
0x00000000c000008c: 0000100f fence.i zero, zero, 0x0
0x00000000c0000090: f1402573 csrr a0, mhartid
0x00000000c0000094: 00100593 li a1, 1
0x00000000c0000098: 00b57063 bgeu a0, a1, 0xc0000098
0x00000000c000009c: 00000413 li s0, 0
0x00000000c00000a0: 00040067 jalr zero, s0
0x0000000000000000: 0000b197 auipc gp, 0xb <_start+0>
0x0000000000000002: 4801b197 auipc gp, 0x4801b <_start+2>
0x0000000000000002: 4801b197 auipc gp, 0x4801b <_start+2>
0x0000000000000002: 4801b197 auipc gp, 0x4801b <_start+2>
0x0000000000000002: 4801b197 auipc gp, 0x4801b <_start+2>
0x0000000000000002: 4801b197 auipc gp, 0x4801b <_start+2>
0x0000000000000002: 4801b197 auipc gp, 0x4801b <_start+2>
grmon4> cont
Error mode (5, Load access fault)
0x0000000000000002: 4801b197 auipc gp, 0x4801b <_start+2>
SIGHUP
After further debugging, we found that the jump to the ROM address is actually the result of an exception rather than its cause.
Immediately after executing the first instruction, we observe:
grmon4> step
0x0000000000000000: 0000b197 auipc gp, 0xb <_start+0>
grmon4> reg mcause mtval mepc
mcause = 5 (0x0000000000000005)
mtval = 2147483664 (0x0000000080000010)
mepc = 0 (0x0000000000000000)
-
To move from the provided reference design to our previous working setup, we replaced
ahb2axi_mig4_7serieswithaxi_mig4_7series. With this change, execution no longer jumps to the ROM address, but instead appears to hang inside__bcc_con_outbyte.Interestingly, the behavior depends on the workload:
- Small matrix sizes complete successfully.
- Larger matrices consistently hang.
- Some intermediate sizes occasionally complete, although not for the full benchmark (which performs multiple runs).
This exact software binary executes correctly on our 2024.4 design. The execution trace is shown below:
grmon4> step 5
0x0000000000000000: 0000b197 auipc gp, 0xb <_start+0>
0x0000000000000004: 48018193 addi gp, gp, 1152 <_start+4>
0x0000000000000008: 6299 lui t0, 0x6 <_start+8>
0x000000000000000a: 3002a073 csrs mstatus, t0 <_start+10>
0x000000000000000e: 00301073 csrw fcsr, zero <_start+14>
grmon4> cont
** DEVICE INFORMATION **
- CPU: NOEL-V (64-bits)
> Optimized implementation
** TEST INFORMATION **
Matrix Multiplication:
- Input matrix A (8-bit): 32 x 32
- Input matrix B (8-bit): 32 x 32
- Output matrix (32-bit): 32 x 32
** TEST EXECUTION **
Interrupted!
0x0000000000009f8c: 2007f793 andi a5, a5, 512 <__bcc_con_outbyte+62>
SIGINT
Has anyone experienced similar issues when using the new noelv-xilinx-vcu118 design in the 2026.2 release, or have any suggestions on what might be causing this behavior? Any pointers or debugging ideas would be greatly appreciated.
Thank you in advance for your help.
Marc