Skip to content

RDMA/mlx5: Use workqueue polling with PREEMPT_RT - #301

Merged
gratian merged 1 commit into
ni:nilrt/26.8/6.18from
gratian:dev/nilrt/26.8/6.18
Sep 29, 2026
Merged

gratian merged 1 commit into
ni:nilrt/26.8/6.18from
gratian:dev/nilrt/26.8/6.18

Conversation

@gratian

@gratian gratian commented Sep 29, 2026

Copy link
Copy Markdown

On PREEMPT_RT kernels the mlx5 completion interrupt thread crashes with a general protection fault in irq_poll_complete():

Oops: general protection fault, probably for non-canonical address 0xdead000000000108: 0000 [#1] SMP NOPTI
CPU: 0 UID: 0 PID: 731 Comm: irq/234-mlx5_co Tainted: P S U O 6.18.52-rt6 #1 PREEMPT_{RT,(full)}
RIP: 0010:irq_poll_complete+0xe/0x40
RAX: dead000000000122 RBX: ffff89ed90492c50 RCX: 0000000000000216
RDX: dead000000000100 RSI: 0000000000000087 RDI: ffff89ed90492c50
Call Trace:

ib_poll_handler+0x5e/0xd0 [ib_core]
irq_poll_softirq+0x9f/0x110
handle_softirqs.isra.0+0xbe/0x270
__local_bh_enable_ip+0x6f/0xb0
irq_forced_thread_fn+0x41/0x50
irq_thread+0x18c/0x270
kthread+0xfe/0x210
ret_from_fork+0x178/0x1e0
ret_from_fork_asm+0x1a/0x30

The fault is in the list_del() inlined from __irq_poll_complete(). Both the next and prev pointers of the irq_poll list node already hold LIST_POISON1 (RDX register) and LIST_POISON2 (RAX register), so the node was unlinked once before this call. The CQ was not freed and reused, because freed memory would not still hold the exact list poison values.

The root cause is irq_poll code uses IRQ_POLL_F_SCHED to mean both "queued on the per-CPU list" and "owned by the poller", but it updates the flag and the list membership in separate steps. The per-CPU list is protected only by local_irq_save(). ib_poll_handler() also gives up ownership with irq_poll_complete() and then calls irq_poll_sched() again after re-arming the CQ. On PREEMPT_RT, irq_forced_thread_fn() runs the completion handler and the softirq that follows it in a preemptible task, not in hard interrupt context. So the assumption that local_irq_save() is enough for mutual exclusion no longer holds, and the gaps in the IRQ_POLL_F_SCHED state machine can be hit.

The only IB_POLL_SOFTIRQ CQs that exist in this configuration are the ones mlx5_ib allocates internally for UMR and for the GSI QP. Both use completion vector 0, which matches the crashing irq/234-mlx5_co thread on CPU 0.

If PREEMPT_RT is enabled, allocate both CQs with IB_POLL_WORKQUEUE. Completions are then handled by ib_comp_wq, which is a WQ_HIGHPRI bound workqueue, and irq_poll is no longer used for these CQs. Neither CQ is on a data path where latency matters. UMR submitters already sleep in wait_for_completion(), and GSI carries MAD/CM traffic. Handling their completions in a preemptible, schedulable context also fits PREEMPT_RT better than softirq processing.

Assisted-by: Copilot:Claude Opus 5.5
Signed-off-by: Gratian Crisan gratian.crisan@emerson.com
(cherry picked from commit e452044)

On PREEMPT_RT kernels the mlx5 completion interrupt thread crashes with
a general protection fault in irq_poll_complete():

  Oops: general protection fault, probably for non-canonical address 0xdead000000000108: 0000 [#1] SMP NOPTI
  CPU: 0 UID: 0 PID: 731 Comm: irq/234-mlx5_co Tainted: P S   U     O        6.18.52-rt6 #1 PREEMPT_{RT,(full)}
  RIP: 0010:irq_poll_complete+0xe/0x40
  RAX: dead000000000122 RBX: ffff89ed90492c50 RCX: 0000000000000216
  RDX: dead000000000100 RSI: 0000000000000087 RDI: ffff89ed90492c50
  Call Trace:
   <TASK>
   ib_poll_handler+0x5e/0xd0 [ib_core]
   irq_poll_softirq+0x9f/0x110
   handle_softirqs.isra.0+0xbe/0x270
   __local_bh_enable_ip+0x6f/0xb0
   irq_forced_thread_fn+0x41/0x50
   irq_thread+0x18c/0x270
   kthread+0xfe/0x210
   ret_from_fork+0x178/0x1e0
   ret_from_fork_asm+0x1a/0x30
   </TASK>

The fault is in the list_del() inlined from __irq_poll_complete().
Both the next and prev pointers of the irq_poll list node already hold
LIST_POISON1 (RDX register) and LIST_POISON2 (RAX register), so the
node was unlinked once before this call. The CQ was not freed and
reused, because freed memory would not still hold the exact list
poison values.

The root cause is irq_poll code uses IRQ_POLL_F_SCHED to mean both
"queued on the per-CPU list" and "owned by the poller", but it updates
the flag and the list membership in separate steps. The per-CPU list is
protected only by local_irq_save(). ib_poll_handler() also gives up
ownership with irq_poll_complete() and then calls irq_poll_sched() again
after re-arming the CQ. On PREEMPT_RT, irq_forced_thread_fn() runs the
completion handler and the softirq that follows it in a preemptible
task, not in hard interrupt context. So the assumption that
local_irq_save() is enough for mutual exclusion no longer holds, and the
gaps in the IRQ_POLL_F_SCHED state machine can be hit.

The only IB_POLL_SOFTIRQ CQs that exist in this configuration are the
ones mlx5_ib allocates internally for UMR and for the GSI QP. Both use
completion vector 0, which matches the crashing irq/234-mlx5_co thread
on CPU 0.

If PREEMPT_RT is enabled, allocate both CQs with IB_POLL_WORKQUEUE.
Completions are then handled by ib_comp_wq, which is a WQ_HIGHPRI bound
workqueue, and irq_poll is no longer used for these CQs. Neither CQ is
on a data path where latency matters. UMR submitters already sleep in
wait_for_completion(), and GSI carries MAD/CM traffic. Handling their
completions in a preemptible, schedulable context also fits PREEMPT_RT
better than softirq processing.

Assisted-by: Copilot:Claude Opus 5.5
Signed-off-by: Gratian Crisan <gratian.crisan@emerson.com>
(cherry picked from commit e452044)
@gratian gratian self-assigned this Sep 29, 2026
@gratian
gratian merged commit 66ab95f into ni:nilrt/26.8/6.18 Sep 29, 2026
1 check failed
@gratian

gratian commented Sep 29, 2026

Copy link
Copy Markdown
Author

This is a straight cherry-pick of already approved PR #300 into the 26.8 release branch w/o conflicts. Bypassing review.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant