RDMA/mlx5: Use workqueue polling with PREEMPT_RT - #301
Merged
Merged
Conversation
On PREEMPT_RT kernels the mlx5 completion interrupt thread crashes with a general protection fault in irq_poll_complete(): Oops: general protection fault, probably for non-canonical address 0xdead000000000108: 0000 [#1] SMP NOPTI CPU: 0 UID: 0 PID: 731 Comm: irq/234-mlx5_co Tainted: P S U O 6.18.52-rt6 #1 PREEMPT_{RT,(full)} RIP: 0010:irq_poll_complete+0xe/0x40 RAX: dead000000000122 RBX: ffff89ed90492c50 RCX: 0000000000000216 RDX: dead000000000100 RSI: 0000000000000087 RDI: ffff89ed90492c50 Call Trace: <TASK> ib_poll_handler+0x5e/0xd0 [ib_core] irq_poll_softirq+0x9f/0x110 handle_softirqs.isra.0+0xbe/0x270 __local_bh_enable_ip+0x6f/0xb0 irq_forced_thread_fn+0x41/0x50 irq_thread+0x18c/0x270 kthread+0xfe/0x210 ret_from_fork+0x178/0x1e0 ret_from_fork_asm+0x1a/0x30 </TASK> The fault is in the list_del() inlined from __irq_poll_complete(). Both the next and prev pointers of the irq_poll list node already hold LIST_POISON1 (RDX register) and LIST_POISON2 (RAX register), so the node was unlinked once before this call. The CQ was not freed and reused, because freed memory would not still hold the exact list poison values. The root cause is irq_poll code uses IRQ_POLL_F_SCHED to mean both "queued on the per-CPU list" and "owned by the poller", but it updates the flag and the list membership in separate steps. The per-CPU list is protected only by local_irq_save(). ib_poll_handler() also gives up ownership with irq_poll_complete() and then calls irq_poll_sched() again after re-arming the CQ. On PREEMPT_RT, irq_forced_thread_fn() runs the completion handler and the softirq that follows it in a preemptible task, not in hard interrupt context. So the assumption that local_irq_save() is enough for mutual exclusion no longer holds, and the gaps in the IRQ_POLL_F_SCHED state machine can be hit. The only IB_POLL_SOFTIRQ CQs that exist in this configuration are the ones mlx5_ib allocates internally for UMR and for the GSI QP. Both use completion vector 0, which matches the crashing irq/234-mlx5_co thread on CPU 0. If PREEMPT_RT is enabled, allocate both CQs with IB_POLL_WORKQUEUE. Completions are then handled by ib_comp_wq, which is a WQ_HIGHPRI bound workqueue, and irq_poll is no longer used for these CQs. Neither CQ is on a data path where latency matters. UMR submitters already sleep in wait_for_completion(), and GSI carries MAD/CM traffic. Handling their completions in a preemptible, schedulable context also fits PREEMPT_RT better than softirq processing. Assisted-by: Copilot:Claude Opus 5.5 Signed-off-by: Gratian Crisan <gratian.crisan@emerson.com> (cherry picked from commit e452044)
Author
|
This is a straight cherry-pick of already approved PR #300 into the 26.8 release branch w/o conflicts. Bypassing review. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
On PREEMPT_RT kernels the mlx5 completion interrupt thread crashes with a general protection fault in irq_poll_complete():
Oops: general protection fault, probably for non-canonical address 0xdead000000000108: 0000 [#1] SMP NOPTI
CPU: 0 UID: 0 PID: 731 Comm: irq/234-mlx5_co Tainted: P S U O 6.18.52-rt6 #1 PREEMPT_{RT,(full)}
RIP: 0010:irq_poll_complete+0xe/0x40
RAX: dead000000000122 RBX: ffff89ed90492c50 RCX: 0000000000000216
RDX: dead000000000100 RSI: 0000000000000087 RDI: ffff89ed90492c50
Call Trace:
ib_poll_handler+0x5e/0xd0 [ib_core]
irq_poll_softirq+0x9f/0x110
handle_softirqs.isra.0+0xbe/0x270
__local_bh_enable_ip+0x6f/0xb0
irq_forced_thread_fn+0x41/0x50
irq_thread+0x18c/0x270
kthread+0xfe/0x210
ret_from_fork+0x178/0x1e0
ret_from_fork_asm+0x1a/0x30
The fault is in the list_del() inlined from __irq_poll_complete(). Both the next and prev pointers of the irq_poll list node already hold LIST_POISON1 (RDX register) and LIST_POISON2 (RAX register), so the node was unlinked once before this call. The CQ was not freed and reused, because freed memory would not still hold the exact list poison values.
The root cause is irq_poll code uses IRQ_POLL_F_SCHED to mean both "queued on the per-CPU list" and "owned by the poller", but it updates the flag and the list membership in separate steps. The per-CPU list is protected only by local_irq_save(). ib_poll_handler() also gives up ownership with irq_poll_complete() and then calls irq_poll_sched() again after re-arming the CQ. On PREEMPT_RT, irq_forced_thread_fn() runs the completion handler and the softirq that follows it in a preemptible task, not in hard interrupt context. So the assumption that local_irq_save() is enough for mutual exclusion no longer holds, and the gaps in the IRQ_POLL_F_SCHED state machine can be hit.
The only IB_POLL_SOFTIRQ CQs that exist in this configuration are the ones mlx5_ib allocates internally for UMR and for the GSI QP. Both use completion vector 0, which matches the crashing irq/234-mlx5_co thread on CPU 0.
If PREEMPT_RT is enabled, allocate both CQs with IB_POLL_WORKQUEUE. Completions are then handled by ib_comp_wq, which is a WQ_HIGHPRI bound workqueue, and irq_poll is no longer used for these CQs. Neither CQ is on a data path where latency matters. UMR submitters already sleep in wait_for_completion(), and GSI carries MAD/CM traffic. Handling their completions in a preemptible, schedulable context also fits PREEMPT_RT better than softirq processing.
Assisted-by: Copilot:Claude Opus 5.5
Signed-off-by: Gratian Crisan gratian.crisan@emerson.com
(cherry picked from commit e452044)