Re: Seeking Assistance with Spin Lock Usage and Resolving Hard LOCKUP Error
From: Muni Sekhar <munisekharrms@gmail.com>
Here is a brief overview of how I have implemented spin locks in my module:
spinlock_t my_spinlock; // Declare a spin lock
// In ISR context (interrupt handler): spin_lock_irqsave(&my_spinlock, flags); // ... Critical section ... spin_unlock_irqrestore(&my_spinlock, flags);
// In process context: (struct file_operations.read) spin_lock(&my_spinlock); // ... Critical section ... spin_unlock(&my_spinlock);
from my understanding, you have the usage backwards. It is the irqsave/irqrestore versions that should be used within process context to prevent the interrupt from being handled on the same cpu while executing in your critical section. The use of irqsave/irqrestore within the isr itself is ok, although perhaps unnecessary. It depends on whether the interrupt can occur again while you are servicing the interrupt (whether on this cpu or another). Usually (?) the same interrupt does not nest, unless you have explicitly coded to allow it (for example, by acknowledging and re-enabling the interrupt early in your ISR). Certainly the spinlock is necessary to protect the critical section from running in an isr on one cpu and process space on another cpu.
From a lockup perspective, not doing the irqsave/irqrestore from process context could explain your problem. Also look for code (anywhere!) that blindly enables interrupts, rather than doing irqrestore from a prior irqsave.
On Thu, May 9, 2024 at 8:56 PM Billie Alsup (balsup) <balsup@cisco.com> wrote:
From: Muni Sekhar <munisekharrms@gmail.com>
Here is a brief overview of how I have implemented spin locks in my module:
spinlock_t my_spinlock; // Declare a spin lock
// In ISR context (interrupt handler): spin_lock_irqsave(&my_spinlock, flags); // ... Critical section ... spin_unlock_irqrestore(&my_spinlock, flags);
// In process context: (struct file_operations.read) spin_lock(&my_spinlock); // ... Critical section ... spin_unlock(&my_spinlock);
from my understanding, you have the usage backwards. It is the irqsave/irqrestore versions that should be used within process context to prevent the interrupt from being handled on the same cpu while executing in your critical section.
The use of irqsave/irqrestore within the isr itself is ok, although perhaps unnecessary. It depends on whether the interrupt can occur again while you are servicing the interrupt (whether on this cpu or another). Usually (?) the same interrupt does not nest, unless you have explicitly coded to allow it (for example, by acknowledging and re-enabling the interrupt early in your ISR). Certainly the spinlock is necessary to protect the critical section from running in an isr on one cpu and process space on another cpu.
In the scenario where an interrupt occurs while we are servicing the interrupt, and in the scenario where it doesn't occur while we are servicing the interrupt, when should we use the spin_lock_irqsave/spin_unlock_irqrestore APIs?
From a lockup perspective, not doing the irqsave/irqrestore from process context could explain your problem. Also look for code (anywhere!) that blindly enables interrupts, rather than doing irqrestore from a prior irqsave.
-- Thanks, Sekhar
From: Muni Sekhar <munisekharrms@gmail.com> In the scenario where an interrupt occurs while we are servicing the interrupt, and in the scenario where it doesn't occur while we are servicing the interrupt, when should we use the spin_lock_irqsave/spin_unlock_irqrestore APIs?
In my experience, the interrupts are masked by the infrastructure before invoking the interrupt service routine. So unless you explicitly re-enable them, there shouldn't be a nested interrupt for the same interrupt number. It is the code run at process context that must be protected using the irqsave/irqrestore versions. You want to not only enter the critical section, but also prevent the interrupt from occurring (on the same cpu at least). If you enter the critical section in process context, but then take an interrupt and attempt to again enter the critical section, then your interrupt routine will deadlock. the interrupt routine will never be able to acquire the lock, and the process context code that was interrupted will never be able to complete to release the lock. So the process context code requires the irqsave/irqrestore variant to not only take the lock, but also prevent a competing interrupt routine from being triggered while you hold the lock. Bottom line is that if a critical section can be entered via both process context and interrupt context, then the process context invocation should use the irqsave/irqrestore variants to disable the interrupt before taking the lock. If it is common code shared between process context and interrupt context, then there is no harm in calling the irqsave/irqrestore version from both contexts. Otherwise, the standard spin_lock/spin_unlock variants (without irqsave/irqrestore) would be used for a critical section shared by multiple threads (different cpus), or when your code has already (separately) handled disabling interrupts as needed before invoking spin_lock.
On Thu, May 9, 2024 at 11:27 PM Billie Alsup (balsup) <balsup@cisco.com> wrote:
From: Muni Sekhar <munisekharrms@gmail.com> In the scenario where an interrupt occurs while we are servicing the interrupt, and in the scenario where it doesn't occur while we are servicing the interrupt, when should we use the spin_lock_irqsave/spin_unlock_irqrestore APIs?
In my experience, the interrupts are masked by the infrastructure before invoking the interrupt service routine. So unless you explicitly re-enable them, there shouldn't be a nested interrupt for the same interrupt number.
It is the code run at process context that must be protected using the irqsave/irqrestore versions. You want to not only enter the critical section, but also prevent the interrupt from occurring (on the same cpu at least). If you enter the critical section in process context, but then take an interrupt and attempt to again enter the critical section, then your interrupt routine will deadlock. the interrupt routine will never be able to acquire the lock, and the process context code that was interrupted will never be able to complete to release the lock. So the process context code requires the irqsave/irqrestore variant to not only take the lock, but also prevent a competing interrupt routine from being triggered while you hold the lock.
Bottom line is that if a critical section can be entered via both process context and interrupt context, then the process context invocation should use the irqsave/irqrestore variants to disable the interrupt before taking the lock. If it is common code shared between process context and interrupt context, then there is no harm in calling the irqsave/irqrestore version from both contexts.
Thanks a lot for the detailed clarification.
Otherwise, the standard spin_lock/spin_unlock variants (without irqsave/irqrestore) would be used for a critical section shared by multiple threads (different cpus), or when your code has already (separately) handled disabling interrupts as needed before invoking spin_lock.
-- Thanks, Sekhar
participants (2)
-
Billie Alsup (balsup) -
Muni Sekhar