| CVE |
Vendors |
Products |
Updated |
CVSS v3.1 |
| In the Linux kernel, the following vulnerability has been resolved:
net: wwan: t7xx: validate the netif index in t7xx_ccmni_recv_skb()
The netif index carried in the DPMAIF PIT header is five bits wide,
but ccmni_inst[] only has room for NIC_DEV_MAX (21) entries.
t7xx_ccmni_recv_skb() indexes the array without a bounds check, so
indexes 21 to 31 read past it. The out-of-bounds value lands in the
callback table that follows the array, which is never NULL, so the
existing !ccmni check does not catch it and the driver dereferences
whatever sits there as a struct t7xx_ccmni.
Drop the skb when the index is out of range.
Verified in a QEMU guest with a fault injector setting the netif
index to 25: the unpatched driver reads a value past ccmni_inst[],
which lands in the callback table, and dereferences it far enough to
queue the skb. With this check the packet is dropped. Well-formed
traffic on index 0 is unaffected.
Changes in v2: none. |
| In the Linux kernel, the following vulnerability has been resolved:
net: wwan: mhi_wwan_mbim: guard against a cyclic NDP chain
The NDP traversal in mhi_mbim_rx() only stops when wNextNdpIndex is
zero. Nothing requires the offsets to advance, so a modem that
points an NDP at itself, or at an earlier NDP, keeps the loop
spinning forever on one CPU.
Break out when the next NDP offset is not larger than the current
one.
Verified in a QEMU guest with a fault injector feeding the driver's
receive callback an NTB whose single NDP points at itself: the
unpatched driver spins in mhi_mbim_rx() with one CPU pinned at 100%
and the thread never returns. With this check the loop terminates
within one iteration.
Changes in v2: move the non-increasing check to the wNextNdpIndex
retrieval site, as suggested by Loic Poulain, instead of tracking
the previous offset in a separate variable. |
| In the Linux kernel, the following vulnerability has been resolved:
net: wwan: mhi_wwan_mbim: check skb_copy_bits() return value
mhi_mbim_rx() ignores the return value of skb_copy_bits() when it
copies each datagram out of the NTB. The datagram offset and length
come from the DPE, which is only checked to lie within the NTB
itself, so a modem can point a datagram outside the received skb.
The copy then fails and the freshly allocated skbn is passed to
netif_rx() with its uninitialized contents still in place, leaking
kernel heap memory into the network stack.
Free the skb and account an error when the copy fails.
Verified in a QEMU guest with a fault injector pointing a DPE
outside the received NTB: the copy fails, and the unpatched driver
hands the uninitialized skbn to the network stack (observed as
"unknown protocol" on bytes that were never written). With this
check the failed datagram is dropped and counted as an rx error.
Changes in v2: factor the free-and-count sequence out into
mhi_mbim_rx_drop(), shared with the unknown-protocol path, as
suggested by Loic Poulain. |
| In the Linux kernel, the following vulnerability has been resolved:
net/sched: act_api: release tail references on DELACTION failure
A batched RTM_DELACTION request takes a temporary reference on each
action before attempting any deletion. tcf_action_delete() clears
each processed slot and drops its temporary reference before attempting
the deletion. If deletion fails, tca_action_gd() calls
tcf_action_put_many() to release the remaining references, but its
tcf_act_for_each_action() iterator stops at the first NULL slot.
When a batch stops at an action bound to a filter, this leaks a
reference on each subsequent action. A later delete of an unbound
action can then return success without removing it from the IDR.
Walk the full array in tcf_action_put_many() and skip NULL slots to
release the references held on the unprocessed actions. |
| In the Linux kernel, the following vulnerability has been resolved:
net/sched: hhf: cap hh_flows_limit at change time
hhf_change() stores TCA_HHF_HH_FLOWS_LIMIT with no upper bound. A huge
hh_flows_limit lets each new heavy-hitter flow pass the
hh_flows_current_cnt check in alloc_new_hh() and forces a fixed-size
kzalloc(GFP_ATOMIC) per flow under spoofed traffic, for unbounded memory
growth.
Bound the attribute with NLA_POLICY_MAX() at 2*HH_FLOWS_CNT (the
hhf_init() default) and report the rejected value via extack. The
deprecated nested parse is kept: legacy tc does not set NLA_F_NESTED on
TCA_OPTIONS. Configs relying on hh_limit above the default were relying
on unbounded, unsafe behaviour and are not supported going forward.
hhf_init() also ran hhf_change() before setting the default
hh_flows_limit, so a user-supplied hh_limit at add time was clobbered
back to 2048. Set the default before hhf_change() so the configured
value sticks.
This is a follow-up to commit eb56a495f59b ("net/sched: hhf: clamp
quantum in change and init paths"), which bounded the quantum of the
same qdisc; the hh_flows_limit bound is the remaining unbounded knob of
that series' scope.
Conditions to recreate the bug: CAP_NET_ADMIN in a user namespace;
tc qdisc change dev X root hhf hh_limit 4294967295 succeeds and the
value is echoed by tc qdisc show, unbounding heavy-hitter flow
allocations; also tc qdisc add dev X root hhf hh_limit 500 stores 2048
instead of 500. |
| In the Linux kernel, the following vulnerability has been resolved:
net/packet: clear RX owner on VNET header error
Commit 61fad6816fc1 ("net/packet: tpacket_rcv: avoid a producer race
condition") added rx_owner_map and made tpacket_rcv() claim a V1 or V2
ring slot before converting the virtio-net header. If the conversion
fails, the drop path leaves the slot claimed.
With a one-frame TPACKET_V2 ring, an unsupported UDP GSO packet leaves
the only slot unavailable, so the ring also drops the next valid packet.
Clear the ownership bit on this error path. TPACKET_V3 already clears
its block state here. |
| In the Linux kernel, the following vulnerability has been resolved:
scsi: core: Validate MODE SENSE lengths in scsi_cdl_enable()
scsi_cdl_enable() uses length fields returned by MODE SENSE to locate
the ATA feature mode page in a 64-byte stack buffer. A target can report
a total length shorter than its mode header and block descriptors. The
unsigned subtraction used for the MODE SELECT length can wrap, and the
separately computed buf_data can point beyond buf.
During automatic scan, enable is false, so the read-modify-write of
buf_data[4] can clear the low two bits of a target-selected
out-of-bounds stack byte. scsi_mode_select() can then copy up to 64
bytes from outside the buffer into the outgoing MODE SELECT payload,
disclosing stack contents to the target.
This is reachable while scanning a USB storage device that identifies as
an ATA device and advertises CDL support. No filesystem mount or
userspace access to the block device is required.
On upstream commit cee9395acd80 ("Linux 7.3-rc1"), a build-specific,
one-vCPU QEMU/Raw Gadget proof using QEMU-only multi-UDC allocator
sampling executed a fixed proof command inside the guest and created a
UID-0-owned marker during automatic enumeration, with KASLR and NX
enabled.
The issue was independently found during security research at Drivesec
S.r.l.
Cap the available length to the buffer size. Validate and consume the
mode header and block descriptor lengths before using the page, and
require the five bytes needed to access the CDL field. |
| In the Linux kernel, the following vulnerability has been resolved:
xfrm: serialize state GC with device state flush
The deferred-device pass in xfrm_dev_state_flush() finds states under
xfrm_state_dev_gc_lock, but drops the lock before calling
xfrm_dev_state_free() because the driver callback may sleep. The device
GC list does not hold an xfrm_state reference, so the state GC worker can
destroy the same state concurrently.
The race can proceed as follows:
CPU 0 CPU 1
find x on the device GC list
drop xfrm_state_dev_gc_lock
read x->xso.dev
xfrm_state_gc_destroy(x)
xfrm_dev_state_free(x)
xfrm_state_free(x)
continue xfrm_dev_state_free(x)
Both paths can invoke the driver callback and drop the device reference.
CPU 0 can also access the xfrm_state after CPU 1 has freed it.
KASAN reported:
BUG: KASAN: slab-use-after-free in xfrm_dev_state_free+0x24c/0x2a0
Read of size 8 at addr ffff88810bbaa960 by task poc/102
Call Trace:
xfrm_dev_state_free+0x24c/0x2a0
xfrm_dev_state_flush+0x353/0x400
xfrm_dev_event+0x26d/0x3a0
notifier_call_chain+0xc0/0x280
__dev_notify_flags+0x169/0x250
netif_change_flags+0xe7/0x160
dev_change_flags+0x96/0x220
devinet_ioctl+0x7f4/0x1880
Allocated by task 87:
xfrm_state_alloc+0x1e/0x5c0
xfrm_add_sa+0xe7f/0x5820
xfrm_user_rcv_msg+0x4f3/0x940
Freed by task 57:
kmem_cache_free+0xcb/0x3d0
xfrm_state_gc_task+0x4a8/0x650
process_one_work+0x63a/0x1070
Serialize xfrm_state destruction against the deferred-device pass with a
mutex. Keep xfrm_state_dev_gc_lock limited to list operations and retain
the existing callback and device-reference release ordering. |
| In the Linux kernel, the following vulnerability has been resolved:
xfrm: use hlist_del_init_rcu for state_cache and state_cache_input
Commit 14acf9652e56 ("xfrm: defensively unhash xfrm_state lists in
__xfrm_state_delete") converted bydst/bysrc/byseq/byspi from
hlist_del_rcu() to hlist_del_init_rcu() so that a second
__xfrm_state_delete() on the same object becomes a no-op rather than a
write through LIST_POISON pprev. It missed state_cache and
state_cache_input, which kept hlist_del_rcu():
- hlist_del_rcu() leaves pprev = LIST_POISON2 (non-NULL), so
hlist_unhashed() returns false.
- hlist_del_init_rcu() leaves pprev = NULL, so hlist_unhashed()
returns true.
A second __xfrm_state_delete() therefore enters __hlist_del() on the
already-deleted state_cache/state_cache_input nodes and does
WRITE_ONCE(*pprev, next) through LIST_POISON2 — a write use-after-free
once the slab is reused. The corruption can in turn cause a subsequent
hlist_for_each_entry_rcu traversal to follow a dangling next pointer,
producing the read use-after-free reported in xfrm_input_state_lookup().
Switch state_cache and state_cache_input to hlist_del_init_rcu() to
match the other four lists, closing the write use-after-free and, with
it, the read use-after-free it spawns. |
| In the Linux kernel, the following vulnerability has been resolved:
xfrm: save input state data before secpath resets
xfrm_input() stores the current xfrm_state in the skb secpath while it
continues receive-side processing. Some input paths can reset that secpath
before xfrm_input() has finished dereferencing the state.
Receive callback users such as VTI and XFRM interfaces can reset the
secpath. The VTI receive path does so before checking whether the packet
crosses network namespaces, while the XFRM interface path does so only for
cross-network-namespace packets. The XFRM_MAX_DEPTH error path can also
reset the secpath before the final drop callback reports the current
state's protocol.
If secpath_reset() drops the last state reference while the state is
concurrently deleted, xfrm_input() can still dereference the freed state
when selecting transport_finish() or reporting the drop callback protocol.
Save the state protocol on the stack while the state is still valid,
and use the already saved address family for transport_finish(). A larval
XFRM_STATE_ACQ state has no type, so retain nexthdr as its protocol. This
preserves the existing drop-path fallback while avoiding the post-reset
state dereferences without adding an extra state reference to every
received packet. |
| In the Linux kernel, the following vulnerability has been resolved:
mips: select CONFIG_WEAK_REORDERING_BEYOND_LLSC from CONFIG_EYEQ
On I6500 CPU cores, lld and scd give no ordering guarantees (same as all
other instructions). To respect the assumption that arch_cmpxchg() is
fully ordered, we must inject sync instructions above and below our
lld/scd loops using the already in place WEAK_REORDERING_BEYOND_LLSC
infrastructure.
Otherwise, bad things can happen:
[ 34.054496] CPU 3 Unable to handle kernel paging request at virtual address 0000000000000000, epc == a80000080838e01c, ra == a80000080838dfc4
[ 34.054559] Oops[#1]:
[ 34.069561] CPU: 3 UID: 0 PID: 170 Comm: pipe_race Not tainted 7.2.0-rc6-01553-gb73c35220968-dirty #103 VOLUNTARY
[ 34.079932] Hardware name: Mobile EyeQ5 MP5 Evaluation board
[ 34.085592] $ 0 : 0000000000000000 0000000000000001 0000000000000000 0000000000000000
[ 34.093616] $ 4 : a800000808ee2618 000000000b7a879d 0000000000001000 0000000000000000
[ 34.101638] $ 8 : 0000000000e3f2c9 0000000000000000 a800000808a2a9f8 0000000000000000
[ 34.109660] $12 : a8000008139ffcd8 ffffffff84080018 a80000080837fae0 7878787878787878
[ 34.117682] $16 : a800000807e82940 0000000000001000 0000000000000000 0000000000000000
[ 34.125704] $20 : a800000802920e00 a8000008139ffdf8 a800000802649400 0000000000e3f2c9
[ 34.133726] $24 : 0000000000000006 00000001200406e0
[ 34.141783] $28 : a8000008139fc000 a8000008139ffd10 0000000000e3f2c8 a80000080838dfc4
[ 34.149837] epc : a80000080838e01c anon_pipe_read+0xd4/0x428
[ 34.155697] ra : a80000080838dfc4 anon_pipe_read+0x7c/0x428
[ 34.161549] Status: 140000e3 KX SX UX KERNEL EXL IE
[ 34.166551] Cause : 40800408 (ExcCode 02)
[ 34.170574] BadVA : 0000000000000000
[ 34.174161] PrId : 0001b028 (MIPS I6500)
[ 34.178183] Process pipe_race (pid: 170, threadinfo=000000005ca35720, task=00000000e1013890, tls=000000014ebbb780)
[ 34.188568] Stack : a800000802649400 0000000000000000 0000000000000000 a8000008139ffdd0
[ 34.196623] 0000000000000fba a800000808ee0000 0000000000000001 a8000008130c3e80
[ 34.204676] a8000008080d1280 a8000008139ffd58 a8000008139ffd58 1dbd2b22ea1dd500
[ 34.212729] a800000802649400 a800000808ee0000 ffffffffffffffea 0000000000000001
[ 34.220783] 0000000000001000 0000000000000000 00000001200ae518 ffffffffffffffff
[ 34.228836] 000000fffbe0e530 a80000080837edf4 000000fffbe0e530 0000000000000000
[ 34.236890] 0000000000000000 0000000000000000 000000014ebb55a0 0000000000001000
[ 34.244943] 0000000000000001 a800000802649400 0000000000000000 0000000000000000
[ 34.252996] 0000000000000000 0000400400000000 0000000000000000 1dbd2b22ea1dd500
[ 34.261049] 00000000140000e3 a800000802649400 a800000802649400 a800000808ee0000
[ 34.269103] ...
[ 34.271568] Call Trace:
[ 34.274026] [<a80000080838e01c>] anon_pipe_read+0xd4/0x428
[ 34.279533] [<a80000080837edf4>] vfs_read+0x25c/0x318
[ 34.284607] [<a80000080837faac>] ksys_read+0x104/0x138
[ 34.289763] [<a80000080802b9cc>] syscall_common+0x44/0x68
[ 34.295187]
[ 34.296689] Code: f84000cf 02209825 de020010 <dc420000> d8400004 02002825 0040f809 02802025 f84000c3
[ 34.306504]
[ 34.308099] ---[ end trace 0000000000000000 ]---
My initial reproducer was the xdp-tools test suite. A standalone
reproducer would be an lld/scd loop that, when the read is reordered by
the CPU, triggers a fault. We can achieve this from userspace by
stressing an anonymous pipe, which uses a mutex. Program used:
// SPDX-License-Identifier: GPL-2.0
// pipe_race.c - reproducer for MIPS LL/SC reordering vs fs/pipe.c
//
// Two userspace processes on an anonymous pipe:
// parent = writer: tight write() loop
// child = reader: tight read() loop
#define _GNU_SOURCE
#include <assert.h>
#include <errno.h>
#include <sched.h>
#include <signal.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/types.h>
#include
---truncated--- |
| In the Linux kernel, the following vulnerability has been resolved:
memstick: ms_block: destroy io_queue workqueue on removal
msb_init_disk() creates the per-card ordered workqueue msb->io_queue with
alloc_ordered_workqueue(). It is torn down with destroy_workqueue() only
on the init error path; msb_remove() never destroys it. msb_stop() merely
flushes the queue, and neither msb_data_clear() nor put_disk() free it. As
a result every card insert/remove cycle leaks the workqueue and its
kworker, exhausting kernel memory over repeated cycles.
Destroy the workqueue in msb_remove() after the disk has been removed and
the queue drained. |
| In the Linux kernel, the following vulnerability has been resolved:
mm, swap: fix SWAP_USAGE_OFFLIST_BIT collision with real usage count
SWAP_USAGE_OFFLIST_BIT is embedded in the si->inuse_pages usage counter,
and is meant to sit above any value that counter can reach. However, it
is defined from BITS_PER_TYPE(atomic_t), so it is bit 30. On a system
with 4 KiB pages the flag collides with the usage count once that count
reaches 4 TiB.
swap_usage_in_pages() masks bit 30 out, so whenever the real count has
that bit set, every caller of it reads 4 TiB low:
* /proc/swaps understates Used by 4 TiB.
* A raw count of exactly 2^30 masks to zero, so try_to_unuse() takes its
"if (!swap_usage_in_pages(si)) goto success;" early exit and swapoff
tears the device down while pages are still swapped out. Nothing in
the rest of swapoff aborts the teardown, so those pages are lost.
Independently of swapoff, the collision also corrupts the counter and the
plist. On a device in normal use, a free that leaves bit 30 set in the
count makes swap_usage_sub() see the flag where there is only count, and
call add_to_avail_list(). It clears the bit with
fetch_and(~SWAP_USAGE_OFFLIST_BIT), leaving the stored count 4 TiB below
the real one, and calls plist_add() on a device that is already listed,
tripping the WARN_ON(!plist_node_empty(node)) in plist_add() and linking
the node a second time.
Change the definition of SWAP_USAGE_OFFLIST_BIT to be based on
atomic_long_t instead. Note that the usage counter field itself is of
this same type, so it is still a valid bit. |
| In the Linux kernel, the following vulnerability has been resolved:
mm/shrinker: fix bogus set_shrinker_bit() with cgroup.memory=nokmem
With cgroup.memory=nokmem, shrinker_memcg_alloc() bails out early and
never allocates an id, so shrinker->id keeps the 0 it got from the
kzalloc() in shrinker_alloc(). __list_lru_init() then copies that 0 into
lru->shrinker_id, where it looks like a valid bit index.
Nothing calls expand_shrinker_info() on nokmem either, so shrinker_nr_max
stays 0 and every memcg ends up with an empty map (map_nr_max == 0).
deferred_split_folio() hands a real memcg to __list_lru_add() regardless
of whether the lru is memcg aware, so the first THP queued in a cgroup
does set_shrinker_bit(memcg, nid, 0) and trips the bounds check:
WARNING: mm/shrinker.c:212 at set_shrinker_bit+0x7d/0x90, CPU#126
Call Trace:
<TASK>
deferred_split_folio+0x18c/0x220
map_anon_folio_pmd_nopf+0xdd/0x130
map_anon_folio_pmd_pf+0x14/0xb0
do_huge_pmd_anonymous_page+0x1a1/0x620
__handle_mm_fault+0xea9/0x10d0
handle_mm_fault+0xe5/0x320
do_user_addr_fault+0x1cc/0x870
exc_page_fault+0x81/0x1b0
asm_exc_page_fault+0x27/0x30
</TASK>
Harmless, the WARN_ON_ONCE() is what keeps the out of bounds unit[] read
from happening, but the id should not look valid in the first place.
Clear it before returning.
Two other spots could paper over this: drop the id in __list_lru_init()
when nokmem turns memcg_aware off, or make deferred_split_folio() pass
NULL like list_lru_add_obj() does. Both leave shrinker->id lying around
for the next caller, so fix it where the id is handed out. |
| In the Linux kernel, the following vulnerability has been resolved:
mm/vma: correctly unaccount on mmap_prepare() failure
__mmap_setup() accounts memory for relevant mappings via:
security_vm_enough_memory_mm()
-> __vm_enough_memory()
-> vm_acct_memory()
If __mmap_setup() fails, this indicates that this accounting did not take
place, and thus it's appropriate for __mmap_region() to jump to
abort_munmap.
However if call_mmap_prepare() fails, it also jumps there and any accounted
memory is not correctly unaccounted.
Fix this by handling each error separately. |
| In the Linux kernel, the following vulnerability has been resolved:
mm: filemap: retain mapped dropbehind folios
Fault-around can map ready dropbehind folios without going through the
normal page-cache lookup that clears dropbehind. A mapping represents a
competing cached user, so retain the folio instead of forcibly unmapping
it when writeback completes.
For a mapped folio, folio_unmap_invalidate() can call
unmap_mapping_folio(), which takes i_mmap_rwsem and may sleep. Retaining
mapped folios avoids this path when folio_end_dropbehind() runs in
non-preemptible task context.
Tal was able to trigger a sleeping-in-atomic warning due to this [1].
Unmapped dropbehind folios continue through the existing invalidation path. |
| In the Linux kernel, the following vulnerability has been resolved:
KEYS: encrypted: fix integer overflow of datablob_len
encrypted_key_alloc() stores datablob_len in a u16. It is computed from
multiple string and payload lengths. If the result exceeds U16_MAX, the
assignment truncates the allocation size. KASAN reports a 32760-byte
slab-out-of-bounds write when __ekey_init() copies the master key
description into the undersized buffer.
The total payload length stored in key->datalen is also a u16. Use
check_add_overflow() to reject values that do not fit either destination,
and use kzalloc_flex() for the flexible-array allocation. |
| Incorrect Privilege Assignment vulnerability in PublishPress PublishPress Capabilities capability-manager-enhanced allows Privilege Escalation.This issue affects PublishPress Capabilities: from n/a through 2.45.0. |
| In the Linux kernel, the following vulnerability has been resolved:
KEYS: trusted: Fix tpm2_load_cmd() boundary check
tpm2_load_cmd() does boundary checks against the ASN.1 size i.e.,
payload->blob_len. Address this by passing the decoded blob size to
tpm2_load_cmd(), and use it for the boundary checks. |
| In the Linux kernel, the following vulnerability has been resolved:
sched_ext: Fix NULL sched deref in kfunc sub-sched error paths
When the root scheduler has sub-scheds attached, the COMPAT kfunc
wrappers scx_bpf_select_cpu_and() and scx_bpf_dsq_insert_vtime() refuse
the call and report to @p's scheduler:
scx_error(scx_task_sched(p), "... must be used");
The wrappers are reachable with tasks that have no scheduler.
scx_bpf_select_cpu_and() is in the select_cpu kfunc group, which
scx_kfunc_context_filter() opens to BPF_PROG_TYPE_SYSCALL programs;
scx_bpf_dsq_insert_vtime() is in the enqueue_dispatch group, which
ops.enqueue() and ops.dispatch() may call with any KF_RCU task -- the
group has no kf_tasks validation, and scx_dsq_insert_preamble() checks
task ownership with scx_task_on_sched() precisely because @p may be an
arbitrary task.
scx_task_sched(p) is p->scx.sched, which is NULL for tasks past
sched_ext_dead() -- which clears it via scx_disable_and_exit_task() on
exit -- and for idle tasks, which the enable paths skip as they are
never scheduled through SCX. It is also an rcu_dereference_protected()
that expects @p's pi_lock or rq lock, which neither wrapper holds.
Passing NULL to scx_error() reaches scx_vexit(), which dereferences
sch->exit_info, oopsing the kernel.
One concrete trigger exercised while developing the fix: a
BPF_PROG_TYPE_SYSCALL program calling the select_cpu_and wrapper on an
exited-but-not-reaped task while a sub-scheduler was attached (its pid
stays findable while the zombie is unreaped; faulting instruction is
the scx_vexit() prologue "mov r15,[rdi+0x398]" with RDI=NULL and 0x398
the offset of sch->exit_info):
sched_ext: BPF scheduler "kfunc_subsched_null" enabled
sched_ext: BPF sub-scheduler "kfunc_subsched_null" enabled
sched_ext: Unassociated program run_select_cpu_ (id 76)
BUG: kernel NULL pointer dereference, address: 0000000000000398
#PF: supervisor read access in kernel mode
#PF: error_code(0x0000) - not-present page
Oops: Oops: 0000 [#1] SMP NOPTI
CPU: 7 UID: 0 PID: 8201 Comm: kfunc_test_runn Tainted: G W
RIP: 0010:scx_vexit+0x25/0xa0
Code: ... <4c> 8b bf 98 03 00 00 ...
CR2: 0000000000000398
Call Trace:
<TASK>
__scx_exit+0x4f/0x70
scx_bpf_select_cpu_and+0xab/0xb0
bpf_prog_430ed61a7b66e03a_run_select_cpu_and+0x9c/0xe7
? __x64_sys_bpf+0x2c/0x40
bpf_prog_test_run_syscall+0x130/0x2f0
__sys_bpf+0x930/0x10d0
? __x64_sys_bpf+0x2c/0x40
__x64_sys_bpf+0x2c/0x40
do_syscall_64+0xbc/0x460
entry_SYSCALL_64_after_hwframe+0x76/0x7e
</TASK>
Read @p's scheduler under RCU instead, which the wrappers can do from
their guard(rcu)(): fault it when it can be determined, and when it
can't be determined -- @p is a task past sched_ext_dead() or an idle
task -- there is nothing obviously wrong to report, so just refuse the
call as before without faulting any scheduler.
These COMPAT wrappers are scheduled for eventual removal once the
deprecation grace period elapses, but until then -- and regardless of
their removal timeline -- they must not oops the kernel on a task they
are handed. |