| CVE |
Vendors |
Products |
Updated |
CVSS v3.1 |
| In the Linux kernel, the following vulnerability has been resolved:
KVM: PPC: Book3S HV: fix use-after-free in kvmhv_emulate_tlbie_all_lpid()
kvmhv_emulate_tlbie_all_lpid() iterates the nested-guest IDR and drops
mmu_lock before calling kvmhv_emulate_tlbie_lpid(), but does not hold a
reference on the kvm_nested_guest pointer obtained from the IDR. A
concurrent vCPU issuing a single-LPID tlbie (is=2, ric=2) can race
through kvmhv_flush_nested() -> kvmhv_remove_nested() -> idr_remove /
--refcnt -> kvmhv_release_nested() -> kfree(gp) in that window, leaving
the iterating vCPU with a dangling pointer. The subsequent
mutex_lock(&gp->tlb_lock) and accesses to gp->shadow_pgtable,
gp->shadow_lpid and gp->l1_host all touch freed memory. The free path
is fully L1-controlled.
Fix this by incrementing gp->refcnt inside the loop before dropping
mmu_lock, mirroring what kvmhv_get_nested() does, and releasing the
reference with kvmhv_put_nested() after the per-guest work completes.
This is the same get/put discipline already used at every other
call site that drops mmu_lock while holding a nested-guest pointer. |
| In the Linux kernel, the following vulnerability has been resolved:
btrfs: handle lack of space when cleaning up verity items
When enable_verity() hits the qgroup limit, rollback_verity() needs its
own metadata reservation. When the qgroup limit or lack of space refuses
the rollback, the whole filesystem is forced read-only even though the
qgroup limit was for one subvolume only. Also orphan cleanup at the next
mount fails the same way, so the leftover items are never removed: with
-EDQUOT the subvolume stays unreachable, and with -ENOSPC on a full
filesystem the next read-write mount fails.
Start transactions with btrfs_start_transaction_fallback_global_rsv() in
btrfs_orphan_cleanup(), drop_verity_items() and rollback_verity(). Those
calls only delete items and free the space in the end, so they may use
the global reserve and skip the qgroup limit, which avoids -ENOSPC and
-EDQUOT. |
| In the Linux kernel, the following vulnerability has been resolved:
net: lock the socket in sock_gettstamp()
sk->sk_flags must only be changed while holding the socket lock,
because sock_set_flag() and sock_reset_flag() use non atomic
operations (__set_bit() and __clear_bit()).
sock_gettstamp() is one of the last places where a bit of sk->sk_flags
is changed from a syscall without owning the socket lock, through
sock_enable_timestamp(sk, SOCK_TIMESTAMP).
sk_set_memalloc() and sk_clear_memalloc() also change sk->sk_flags
without the socket lock, but their callers (nbd, iscsi_tcp, nvme-tcp,
sunrpc, wireguard) need a careful audit, this will be addressed in a
separate patch.
Jungwoo Lee and Wongi Lee reported an UDP socket use-after-free
caused by this bug: a SIOCGSTAMPNS_NEW ioctl racing with bind()
can cancel the SOCK_RCU_FREE bit that udp_lib_get_port() just set,
because both threads perform a read-modify-write on the same word.
CPU 0 (bind) CPU 1 (SIOCGSTAMPNS_NEW)
-------------------------------- ----------------------------
read sk_flags = F read sk_flags = F
compute F | BIT(SOCK_RCU_FREE) compute F | BIT(SOCK_TIMESTAMP)
store F | BIT(SOCK_RCU_FREE)
sk_add_node_rcu(sk, ...)
store F | BIT(SOCK_TIMESTAMP)
After the lost update, SOCK_RCU_FREE is clear while the socket is
visible to lockless UDP receive lookups. sk_destruct() then frees
the socket immediately instead of waiting for a RCU grace period,
while the receive path still holds a reference-less pointer to it:
BUG: KASAN: slab-use-after-free in ipv4_pktinfo_prepare+0x30/0x410
Read of size 8 at addr ffff888008806610 by task exploit/207
CPU: 0 UID: 1000 PID: 207 Comm: exploit Not tainted 6.12.95+ #1
ipv4_pktinfo_prepare+0x30/0x410
udp_queue_rcv_one_skb+0x51c/0x1180
udp_unicast_rcv_skb+0x109/0x350
ip_protocol_deliver_rcu+0x14b/0x310
ip_local_deliver_finish+0x29d/0x390
ip_local_deliver+0x24d/0x2a0
Only grab the socket lock when SOCK_TIMESTAMP has to be set,
to keep the common case lockless. |
| In the Linux kernel, the following vulnerability has been resolved:
net: ethernet: cortina: Ack RX overrun interrupt correctly
The RX overrun interrupt is reported in interrupt status register 4, but
gmac_irq() acknowledges it using the RX descriptor error bit from status
register 0. For GMAC0 this writes the GMAC1 overrun bit, while for GMAC1
the shift leaves no bit in the 32-bit register.
Acknowledge the same per-port RX overrun bit that was detected. |
| In the Linux kernel, the following vulnerability has been resolved:
net: psp: avoid conflicts with skb->decrypted and sk_validate_xmit_skb()
PSP conflicts with TLS ULP in its usage of both skb->decrypted and
sk->sk_validate_xmit_skb().
Make PSP mutually exclusive with TLS ULP, the only other user of either
of these. As other users of skb->decrypted come along, they can be added
to sk_has_decrypt_user(). It would make sense to also assert that
sk->sk_validate_xmit_skb() is also NULL in both of these setup paths for
similar future proofing, but the PSP listener/sk_clone() path is still
broken and it could be seen as a regression to not allow rx assoc to run
on a child of a listener socket with PSP tx assoc state.
Include all TCP ULPs in the sk_has_decrypt_user() check, even though TLS
is the only one that conflicts with PSP via the decrypted bit. This is
intentional because PSP was not designed to be used with ULPs. It is
best to close off surface area that may make bugs reachable, until
someone wishes to design and test an actual user of PSP with ULPs. |
| In the Linux kernel, the following vulnerability has been resolved:
x86/kprobes: Fix crash when probing CS CALL instructions
When using eBPF to probe CS CALL instructions within a function,
a crash can be triggered.
The eBPF tool probes offset 257 of the __hrtimer_run_queues()
function:
<__hrtimer_run_queues+249>: nopl 0x0(%rax,%rax,1)
<__hrtimer_run_queues+254>: mov %r14,%rdi
<__hrtimer_run_queues+257>: cs call <__x86_indirect_thunk_r12>
<__hrtimer_run_queues+263>: mov %eax,%r12d
<__hrtimer_run_queues+266>: xchg %ax,%ax
<__hrtimer_run_queues+268>: mov %r13,%rdi
Which triggers this crash:
BUG: unable to handle page fault for address: 00000000000f41c9
#PF: supervisor write access in kernel mode
#PF: error_code(0x0002) - not-present page
PGD 0 P4D 0
Oops: 0002 [#1] SMP NOPTI
CPU: 1 PID: 0 Comm: swapper/1 Kdump: loaded Tainted: P
RIP: 0010:__hrtimer_run_queues+0x106/0x230
Note that __hrtimer_run_queues+0x106 is __hrtimer_run_queues+262, which is
at the 6th byte of the above CS CALL instruction. Since the CS CALL
instruction occupies 6 bytes, the exception occurred in the middle of that
call instruction.
The root cause is that when using eBPF tools to probe in the middle of a
function, a kprobe with INT3 is used as the underlying implementation.
During single-step emulation of the original CALL instruction,
int3_emulate_call() assumes that the probed CALL instruction is 5 bytes
long. However, the actual CS-prefixed CALL instruction occupies 6 bytes,
so it constructs an incorrect exception return address. When the CPU
returns from the kprobe handler, the next instruction to be executed is at
the address of the last byte of that CS CALL instruction. Coincidentally,
starting from that address, the CPU fetches and decodes a completely
different instruction, which ultimately triggers a kernel crash.
Fix the issue by using the actual instruction length obtained from
the instruction decoder when constructing the exception return
address, rather than relying on the hardcoded CALL_INSN_SIZE macro.
[ mingo: Refined the changelog ] |
| In the Linux kernel, the following vulnerability has been resolved:
drm/amdgpu: check ras and obj before dereference
nbio_v7_9_handle_ras_controller_intr_no_bifring() dereferences ras and obj
without checking either for NULL. Both amdgpu_ras_get_context() and
amdgpu_ras_find_obj() can return NULL, e.g. during the window between
adev->nbio.ras being set (early in amdgpu_ras_init(), by design, to
enable the fatal-error interrupt as soon as possible) and the PCIE_BIF
ras object actually being created in RAS late_init. Any interrupt in that
window crashes in hard-IRQ context.
This is analogous to commit d190b459b2a4 ("drm/amdgpu: the warning
dereferencing obj for nbio_v7_4"), which fixed the same issue in the
nbio_v7_4 handler.
Found by Linux Verification Center (linuxtesting.org) with SVACE.
(cherry picked from commit c7071767a50a32ed727cf800ac84372429e3b4b3) |
| In the Linux kernel, the following vulnerability has been resolved:
perf: Fix null pointer access in is_include_guest_event()
A typical module unload occurring event when there is an active perf
connection leads to freeing of the pmu pointer. The call log is something
like:
..
__pmu_detach_event
pmu_detach_event
pmu_detach_events
perf_pmu_unregister
..
__pmu_detach_event() sets event->pmu to null. When the perf connection
finally is closed, the following stack trace is observed:
Oops: general protection fault, kernel NULL pointer dereference
...
RIP: 0010:_free_event+0x3e/0x370
...
Call Trace:
...
perf_event_release_kernel+0x260/0x2d0
perf_release+0x12/0x20
A call to mediated_pmu_unaccount_event() inside _free_event() is the root
cause of this crash. Adding a check inside is_include_guest_event() ensures
we don't accidentally access a null pmu ptr. In addition to this, we will
now call mediated_pmu_unaccount_event() before clearing the pmu ptr so that
nr_include_guest_events counts are maintained correctly. |
| In the Linux kernel, the following vulnerability has been resolved:
9p: Fix v9fs_issue_write() to update i_size and remote_i_size
Fix v9fs_issue_write() to update i_size and remote_i_size to the new size
of the server file if we made it larger, using the start fpos and the count
returned by p9_client_write() to calculate the new minimum file size.
This assumes that if the 9P server makes a short write (say it hits
ENOSPC), a reduced count is returned. |
| In the Linux kernel, the following vulnerability has been resolved:
ALSA: core: Fix potential UAF after asynchronous card release
Usually a sound driver releases the resources assigned to the card via
snd_card_free(), and it synchronizes with the whole release procedure.
However, when the card is released asynchronously via
snd_card_free_when_closed() like USB-audio driver, the situation is
slightly different; although the snd_card_disconnect() call at the
disconnection guarantees that any newer accesses will be gated, the
in-flight tasks might be still accessing to the underlying card->dev
device even after the disconnection, which would cause a
use-after-free in the end, as reported by fuzzers.
For addressing the bug above, this patch takes the refcount of
card->dev at initialization of the card object, and releases at its
destructor. This assures the availability of the card->dev in its
whole lifecycle. |
| In the Linux kernel, the following vulnerability has been resolved:
ALSA: usb-audio: Clamp implicit feedback packet count to URB capacity
data_ep_set_params() allocates each data URB for exactly u->packets
isochronous frames, so urb->iso_frame_desc[] has u->packets slots and
ctx->packets is the driver's only record of that limit. For an implicit
feedback sink, snd_usb_queue_pending_output_urbs() overwrites it with the
sync source's packet count, which is calculated independently from the
capture endpoint's parameters. When that count is larger,
prepare_playback_urb() and prepare_silent_urb() can write
iso_frame_desc[] past the allocation; their existing bounds limit payload
bytes, not the descriptor index.
The reproducer uses a high-speed UAC2 device declaring bInterval 1 for
implicit feedback capture (8 packets) and bInterval 4 for playback
(1 packet). On the first capture completion after the stream starts, it
accesses seven descriptors spanning 112 bytes beyond the one-packet URB:
BUG: KASAN: slab-out-of-bounds in prepare_playback_urb (sound/usb/pcm.c:1560)
Write of size 4 at addr ffff88801e696ad0 by task vhci_rx/178
prepare_playback_urb (sound/usb/pcm.c:1560)
prepare_outbound_urb (sound/usb/endpoint.c:340)
snd_usb_queue_pending_output_urbs (sound/usb/endpoint.c:501)
snd_complete_urb (sound/usb/endpoint.c:1834)
__usb_hcd_giveback_urb (drivers/usb/core/hcd.c:1657)
usb_hcd_giveback_urb (drivers/usb/core/hcd.c:1741)
vhci_rx_loop (drivers/usb/usbip/vhci_rx.c:107)
kthread (kernel/kthread.c:436)
The buggy address belongs to the object at ffff88801e696a00
which belongs to the cache kmalloc-256 of size 256
The buggy address is located 0 bytes to the right of
allocated 208-byte region [ffff88801e696a00, ffff88801e696ad0)
Record the allocated packet count per endpoint and clamp both the adopted
count and the packet-size copy to it. Fold the Format Type II delimiter
into urb_packs before the allocation loop so the recorded limit matches
every URB. |
| In the Linux kernel, the following vulnerability has been resolved:
ALSA: virtio: reset device before deleting virtqueues
virtsnd_remove() and virtsnd_freeze() delete the virtqueues before
resetting the device. del_vqs() frees the vring backing, but does not
provide a generic device quiesce operation. In particular, modern
virtio-pci keeps enabled queues active until the device is reset.
Reset the device before deleting the virtqueues so it can no longer
access the vring memory when that memory is released. This also covers
probe failures after DRIVER_OK, which unwind through virtsnd_remove(). |
| In the Linux kernel, the following vulnerability has been resolved:
cifs: Fix server use-after-free in cifs_chan_skip_or_disable()
When a secondary channel is no longer supported by the server,
cifs_chan_skip_or_disable() drops the channel reference with
cifs_put_tcp_session() and then continues to use the server pointer by
calling cifs_signal_cifsd_for_reconnect() on it and reading its
primary_server pointer. cifs_put_tcp_session() can drop the last
reference of the channel and tear it down, so both the channel and the
primary server (whose reference is also dropped by
cifs_put_tcp_session()) can be freed before they are signaled for
reconnect.
Signal the channel and the primary server and capture the primary
server pointer before dropping the channel reference with
cifs_put_tcp_session(). |
| In the Linux kernel, the following vulnerability has been resolved:
exec: Cleanup POSIX timers right after de_thread()
A per-thread CPU timer holds a reference to the PID of the thread it is
attached to and, while it is armed, its node is queued in that thread's
posix_cputimers. The task is looked up by that PID.
When a non-leader thread exec()s, de_thread() changes which task owns
that PID. pid_task(timer->it.cpu.pid, PIDTYPE_PID) then returns NULL,
but the node is still queued on tsk, which is alive. timer_lock_sighand()
takes a failed lookup to mean that the node is already dequeued, so it
has nothing to undo.
begin_new_exec() calls posix_cpu_timers_exit(me) right after
exec_task_namespaces() and that removes the leftover node, so the state
normally stays invisible. But bprm->point_of_no_return is set before
de_thread(), so if unshare_files(), set_mm_exe_file(), exec_mmap() or
exec_task_namespaces() fails, the task dies before it gets there.
exit_itimers() then frees the k_itimer while its node is still queued,
and reaping tsk later erases that freed node from the rbtree.
In short:
the non-leader thread B the parent
timer_create(CLOCK_THREAD_CPUTIME_ID)
timer_settime()
arm_timer() // the node is queued on B
execve()
de_thread(B)
exchange_tids(B, leader) // B's PID now belongs to the leader
release_task(leader)
__exit_signal(leader)
posix_cpu_timers_exit(leader) // cleans leader's queue, not B's
__unhash_process(leader) // that PID has no task anymore
exec_mmap()
mmap_read_lock_killable(old_mm)
kill(B, SIGKILL)
// -EINTR
get_signal()
do_exit()
exit_itimers()
posix_timer_delete()
posix_cpu_timer_del()
posix_timer_unhash_and_free() // freed while still queued
wait4()
release_task(B)
posix_cpu_timers_exit(B)
cleanup_timerqueue()
timerqueue_del() // use-after-free
Move the POSIX timer cleanup right after de_thread() before any of the
later failure conditions brings the task into do_exit().
[ tglx: Move the cleanup right after de_thread() ] |
| In the Linux kernel, the following vulnerability has been resolved:
fs/dax: check zero or empty entry before converting xarray entry
Calling dax_to_folio() with empty entry causes kernel panic below when
booting a VM with DAX enabled storage.
This patch checks empty entry before calling dax_to_folio() on
dax_associate_entry(), dax_disassociate_entry(), and dax_busy_page().
Commit 98c183a4fccf ("fs/dax: don't disassociate zero page entries") added
guards in the associate and disassociate paths, but the guards still come
after dax_to_folio(), and dax_busy_page() still has the same problem.
[ 0.737679] EXT4-fs (pmem0p1): mounted filesystem 79676804-7c8b-491a-b2a6-9bae3c72af70 ro with ordered data mode. Quota mode: disabled.
[ 0.737891] VFS: Mounted root (ext4 filesystem) readonly on device 259:1.
[ 0.739119] devtmpfs: mounted
[ 0.739476] Freeing unused kernel memory: 1920K
[ 0.740156] Run /sbin/init as init process
[ 0.740229] with arguments:
[ 0.740286] /sbin/init
[ 0.740321] with environment:
[ 0.740369] HOME=/
[ 0.740400] TERM=linux
[ 0.743162] Unable to handle kernel paging request at virtual address fffffdffbf000008
[ 0.743285] Mem abort info:
[ 0.743316] ESR = 0x0000000096000006
[ 0.743371] EC = 0x25: DABT (current EL), IL = 32 bits
[ 0.743444] SET = 0, FnV = 0
[ 0.743489] EA = 0, S1PTW = 0
[ 0.743545] FSC = 0x06: level 2 translation fault
[ 0.743610] Data abort info:
[ 0.743656] ISV = 0, ISS = 0x00000006, ISS2 = 0x00000000
[ 0.743720] CM = 0, WnR = 0, TnD = 0, TagAccess = 0
[ 0.743785] GCS = 0, Overlay = 0, DirtyBit = 0, Xs = 0
[ 0.743848] swapper pgtable: 4k pages, 48-bit VAs, pgdp=00000000b9d17000
[ 0.743931] [fffffdffbf000008] pgd=10000000bfa3d403, p4d=10000000bfa3d403, pud=1000000040bfe403, pmd=0000000000000000
[ 0.744070] Internal error: Oops: 0000000096000006 [#1] SMP
[ 0.748888] CPU: 0 UID: 0 PID: 1 Comm: init Not tainted 6.18.4 #1 NONE
[ 0.749421] pstate: 004000c5 (nzcv daIF +PAN -UAO -TCO -DIT -SSBS BTYPE=--)
[ 0.749969] pc : dax_disassociate_entry.constprop.0+0x20/0x50
[ 0.750444] lr : dax_insert_entry+0xcc/0x408
[ 0.750802] sp : ffff80008000b9e0
[ 0.751083] x29: ffff80008000b9e0 x28: 0000000000000000 x27: 0000000000000000
[ 0.751682] x26: 0000000001963d01 x25: ffff0000004f7d90 x24: 0000000000000000
[ 0.752264] x23: 0000000000000000 x22: ffff80008000bcc8 x21: 0000000000000011
[ 0.752836] x20: ffff80008000ba90 x19: 0000000001963d01 x18: 0000000000000000
[ 0.753407] x17: 0000000000000000 x16: 0000000000000000 x15: 0000000000000000
[ 0.753970] x14: ffffbf3154b9ae70 x13: 0000000000000000 x12: ffffbf3154b9ae70
[ 0.754548] x11: ffffffffffffffff x10: 0000000000000000 x9 : 0000000000000000
[ 0.755122] x8 : 000000000000000d x7 : 000000000000001f x6 : 0000000000000000
[ 0.755707] x5 : 0000000000000000 x4 : 0000000000000000 x3 : fffffdffc0000000
[ 0.756287] x2 : 0000000000000008 x1 : 0000000040000000 x0 : fffffdffbf000000
[ 0.756871] Call trace:
[ 0.757107] dax_disassociate_entry.constprop.0+0x20/0x50 (P)
[ 0.757592] dax_iomap_pte_fault+0x4fc/0x808
[ 0.757951] dax_iomap_fault+0x28/0x30
[ 0.758258] ext4_dax_huge_fault+0x80/0x2dc
[ 0.758594] ext4_dax_fault+0x10/0x3c
[ 0.758892] __do_fault+0x38/0x12c
[ 0.759175] __handle_mm_fault+0x530/0xcf0
[ 0.759518] handle_mm_fault+0xe4/0x230
[ 0.759833] do_page_fault+0x17c/0x4dc
[ 0.760144] do_translation_fault+0x30/0x38
[ 0.760483] do_mem_abort+0x40/0x8c
[ 0.760771] el0_ia+0x4c/0x170
[ 0.761032] el0t_64_sync_handler+0xd8/0xdc
[ 0.761371] el0t_64_sync+0x168/0x16c
[ 0.761677] Code: f9453021 f2dfbfe3 cb813080 8b001860 (f9400401)
[ 0.762168] ---[ end trace 0000000000000000 ]---
[ 0.762550] note: init[1] exited with irqs disabled
[ 0.762631] Kernel panic - not syncing: Attempted to kill init! exitcode=0x0000000b |
| In the Linux kernel, the following vulnerability has been resolved:
signal: Prevent exec() race
Hyunwoo debugged the following KASAN UAF splat:
BUG: KASAN: slab-use-after-free in __send_signal_locked+0xb27/0xba0
Write of size 8 at addr ffff888007ed80c8 by task poc/79
...
Call Trace:
__send_signal_locked+0xb27/0xba0
do_send_sig_info+0xa7/0x160
do_send_specific+0x76/0xa0
__x64_sys_tgkill+0x193/0x270
...
Allocated by task 80:
do_timer_create+0x1a4/0x1030
__x64_sys_timer_create+0x145/0x190
...
Freed by task 12:
kmem_cache_free_bulk+0x1f8/0x4a0
kvfree_rcu_bulk+0x14f/0x1c0
kfree_rcu_work+0x128/0x1a0
...
Last potentially related work creation:
kvfree_call_rcu+0x39/0x390
__flush_itimer_signals+0x211/0x320
flush_itimer_signals+0x47/0x90
begin_new_exec+0xa6b/0x28c0
It turned out that this happens with a non-leader exec() as Hyunwoo
explained:
de_thread() calls exchange_tids() before release_task(leader), so the
struct pid held by a SIGEV_THREAD_ID timer created against the leader's tid
now points to the thread which called execve(). pid_task() returns that
thread and lock_task_sighand() on it succeeds.
If the timer signal is blocked, its sigqueue stays queued on the leader's
task::pending. The next expiry of that timer can then run while
release_task() flushes the queue.
posixtimer_send_sigqueue() checks whether the sigqueue is already queued
with a plain list_empty(), which only reads list_head::next.
list_del_init() is not atomic and INIT_LIST_HEAD() stores list_head::next
before list_head::prev, so the check can pass in between. list_add_tail()
queues the entry on the task::pending of the live thread, and the
list_head::prev store from the flush then overwrites the list_head::prev
link that list_add_tail() has just set.
__flush_itimer_signals() does not undo that either. With list_head::prev
pointing at the entry itself, its list_del_init() only stores the same
values again, so the entry is not removed from the list. It is still there
after the last reference is dropped and the timer is freed by RCU, and the
list_add_tail() of a later tgkill() follows that list_head::prev into the
freed timer.
This problem surfaced with the recent commit which moved the sigqueue flush
out of the sighand lock held region.
Hyonwoo proposed to fix this by using list_del_init_careful(), but that
just papers over the problem. After some disucssions and various attempts
to solve it, Eric pointed out that there is no reason to flush
task::pending late in release_task() and it should be done in
exit_signals() already.
As nothing can collect and deliver signals which are queued in a dying
task's pending queue, there is no reason to delay it further.
But it has to be ensured that no signals can be queued into it after that
point. exit_signals() sets PF_EXITING in task::flags, which can be used as
an indicator for this.
Cure it by:
- Preventing signal queueing for task private signals (PIDTYPE_PID) when
the task has PF_EXITING set in __send_signal_locked() and in
posixtimer_send_sigqueue().
- Protecting the unlocked setting of PF_EXITING in exit_signals() for the
task group empty and the group exit case with sighand lock
- Flushing task::pending signals right there.
Optimize that by moving the whole pending list to an on-stack list head
under sighand lock and free the signals without the lock held.
There has been quite some discussion about the lockless flush and the
non-leader exec case on weakly ordered systems. The problem is that a third
party which tries to send a posix timer signal relies on the PID lookup to
find the target task and that lookup might result in the new leader when
the signal was originaly directed to the old leader. In case that the
signal was queued on the old leader then the lockless flush raised a
concern over the following situation:
old_leader new_leader third party
A: flush_list() // list_del_in
---truncated--- |
| In the Linux kernel, the following vulnerability has been resolved:
swiotlb: use the adjusted address for the highmem page lookup
swiotlb_bounce() reads the page frame number from the slot's recorded
orig_addr, then advances orig_addr by tlb_offset to reach the address
the caller asked about. The highmem branch mixes the two: the offset
within the page comes from the adjusted address, the page from the value
before it.
Once the adjustment crosses a page boundary the pair no longer describes
one location, and the whole copy lands one page below the intended one
for a positive tlb_offset, one above for a negative one. DMA_FROM_DEVICE
writes the device data over the wrong page and leaves the intended one
stale, DMA_TO_DEVICE feeds the device from a page the mapping may not
cover. Partial syncs through dma_sync_single_range_for_*() are what make
tlb_offset non-zero.
The branch test is picked the same way, so a slot recorded in lowmem can
be adjusted into highmem and the lowmem path then hands a highmem
address to phys_to_virt().
Take both from orig_addr once it is final and keep pfn in the branch
that uses it. PhysHighMem() asks the question straight from the address,
as dma-debug already does. |
| In the Linux kernel, the following vulnerability has been resolved:
RDMA/ucma: Serialize join and leave on copy_to_user failure
rdma_join_multicast() queues RoCE work that later reads the ucma_multicast
through event->param.ud.private_data, then list_add()s the CMA multicast
at the head of id_priv->mc_list. rdma_leave_multicast() matches only by
sockaddr and destroys the first hit.
ucma_process_join() used to drop ctx->mutex after a successful join and
retake it only if copy_to_user() failed. Two concurrent JOIN_MCAST calls
with the same address can therefore insert a second CMA entry before the
first thread's leave. leave then cancels the newer work and the older
worker still dereferences the ucma_multicast that the first thread frees.
Keep ctx->mutex held from rdma_join_multicast() through copy_to_user() and,
on -EFAULT, through rdma_leave_multicast() so leave cannot miss this join.
Do not leave if join itself failed: that path never published this address
on mc_list, and a leave-by-addr would destroy an earlier successful join. |
| In the Linux kernel, the following vulnerability has been resolved:
RDMA/core: fix refcount bug in iwpm_get_nlmsg_request()
iwpm_get_nlmsg_request() initializes refcount _after_ list_add_tail()
making it accessible to global list where another CPU can kref_get()
on nlmsg_request causing a refcount "addition on 0" bug. Fix this
by initializing kref _before_ list_add_tail() so refcount for
nlmsg_request can be incremented/decremented normally. In addition,
also initialize every field before list_add_tail(). |
| In the Linux kernel, the following vulnerability has been resolved:
openvswitch: avoid reallocating confirmed conntrack labels
ovs_ct_get_conn_labels() adds the labels extension when a conntrack
entry does not have one. Confirmed conntracks can be read locklessly,
so adding an extension may reallocate and free the extension block
while another CPU accesses it.
Only add the extension for unconfirmed conntracks. A confirmed
conntrack without labels now fails the caller's label operation instead
of reallocating its extension storage. |