Skip to content

Next - #180

Merged
heatd merged 51 commits into
masterfrom
next
Aug 22, 2026
Merged

Next#180
heatd merged 51 commits into
masterfrom
next

Conversation

@heatd

@heatd heatd commented Aug 22, 2026

Copy link
Copy Markdown
Owner

No description provided.

heatd added 30 commits August 22, 2026 20:43
When activating pages, the LRU flag needs to be set back.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Pass the current si_status (should be simply the signal that stopped the
task).

Fixes: d5bbf4b ("fork: Implement proper sys_clone semantics")
Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Inactive works are also attached to a pool_workqueue. Therefore, ->data
needs to be set.

Fixes: 19ef128 ("workqueue: add initial version of workqueues")
Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Disallow the same work running multiple times concurrently, by keeping
track of every worker's executing work, plus a check when assigning
work. If the work that's getting assigned is already running, then that
same worker will run it after finishing.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
This all stemmed from PTY write paths, which recursed very awkwardly
immediately.

Since the workqueue implementation was introduced, replace DPC with it,
and use it for every TTY driver, including serial, VT and PTY.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Instead of checking for egid == 0, it's supposed to check for euid == 0
(if the process is capable user-wise).

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Instead of trying to manually piece together the options string (by
parsing $'s), take the whole thing. Whoever is calling crypt() can just
do that, and crypt() will grok its own options.

This fixes hash calculation on some hash formats different than the
static files used in the repo.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
On every fill_pbuf() call, the kernel dropped cmsgs by setting has_cmsg
to false. This is wrong, and trivially explained by:

sendmsg(Message 0, SCM_RIGHTS)
\- creates new pbf, adds fds
sendmsg(Message 1, SCM_RIGHTS)
\- tries to merge with 0, calls fill_pbuf()
  \- gets 0, not an error (merge did not happen), clears has_cmsg
\- allocates new pbf, no longer has cmsgs, silently dropped

This caused problems in programs doing FD passing like OpenSSH sshd,
which send various FDs in a row, without plain messages in between.

Fix it by reworking this code with a unix_sendmsg_data. The consuming
side (fill_pbuf()) consumes it by clearing msg after init.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Fix various weird oddities and bugs in FD passing partial failure, and
socket destruction without consuming, by properly setting pbf->dtor to
unix_pbf_free. This covers every pbf destruction path.

Also cover any odd edgecases in unix_scm_rights() failing midway by
clamping nfiles.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Logic pushing raw input bytes to the line discipline needs to properly
handle backpressure (when the buffer is full).

Do so by setting a flag, that input consumption will look at to further
wake up tty_ldisc_input() work.

This fixes dropped input in large PTY writes.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
In case something fires, lockdep will stop updating its internal data,
and everyone looking at it will get stale content.

Wrap the whole WARN_ON_ONCE with a debug_locks. If lockdep fires, then
debug_locks will be set to 0.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Erroneously, the dgram recvmsg() logic was only discarding the packet
once it reached the end. This is wrong and does not provide DGRAM
semantics on the read side. A short read should discard the whole
packet, instead of providing more of the same packet on the next
recvmsg().

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Add SOCK_SEQPACKET support by using most of STREAM's
connection-based logic, plus strict DGRAM boundary limits.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
For use in C code.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Not clearing the active flag causes confusion in the LRU, since
it increments inactive_file/anon, and then on free it will decrement
active_file/anon. This is obviously wrong.

Clear the flag.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
swap_free() returned the swap in use, not free. Also, SwapTotal: was
being printed twice, instead of the correct SwapFree: label.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
It's possible the batch was completely added to the LRU, which makes it
so no one grabs a lock. This leads to a NULL deref when doing the final
unlock.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
folio_batch_init() was broken (and using it would make the next
folio_batch_add() store out of bounds). folio_batch_count() was also
broken. And ultimately the folio_batch_add()'s return value was also
broken (although it ended up being fine for the one case that the mm
code cares about, which is "is this folio batch full").

Fix it.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Otherwise, if batchlen != 1 (which currently there isn't any caller that
does so), the result ends up being incorrect (a batch filled with the
exact same struct page).

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Add them to the tail of the LRU, and not the head. Otherwise, the next
reclaim run will find them (which results in not much of a rotation at
all).

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Dirty indirect blocks when doing truncation.

This fixes data corruption on partial truncates (e.g truncates that
don't get rid of the indirect block).

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
It was accidentally jumping to out_release_write, so it released a write
it was never able to grab.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Do tid clearing at exec time as well, before the mm is switched to the
new one. This matches what happens in Linux.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Prepare reclaim for large folios.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Prepare for large folios. As a side effect, delete the page batching
logic (in favour of the folio batch).

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Most of it isn't safe for large folios yet (due to pagecache and swap
reasons).

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
heatd added 21 commits August 22, 2026 20:44
RSS and VIRT need to be given in pages, not in kilobytes.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Always reset compound page metadata on new allocations. This makes sure
that new allocations don't accidentally pick up stale info (and
accidentally make page_compound_head() point into a stale head).

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Properly unlock the page (by doign nr_ios = 0 sooner) if ext2_map_page
fails. Otherwise, the if (nr_ios == 0) check in out_err is false, and
the page is never unlocked.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
If the reader is shut down, it's supposed to get 0 (EOF) instead of
-EPIPE. -EPIPE only applies to writers.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Signal peer_nowr to the peer on disconnects. This provides the proper
EOF semantics on the reader's side when the peer disconnects.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Instead of relying on userspace to do it, do it in the kernel. This
makes it so init= is actually useful (and can use e.g /bin/bash
directly).

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Sometimes, when the stars align, the following hang early in boot can be
observed:

atomic<unsigned long>::load (this=0xffffffff81c801b0, order=mem_order::seq_cst) at include/onyx/atomic.hpp:51
atomic<unsigned long>::operator unsigned long (this=0xffffffff81c801b0) at include/onyx/atomic.hpp:175
smp::internal::sync_call_cntrlblk::wait (this=0xffffffff81c801a0, local=<optimized out>, context=<optimized out>) at kernel/smp.cpp:121
0xffffffff8103f144 in smp::sync_call_with_local (f=f@entry=0xffffffff81103b10 <x86_invalidate_tlb(void*)>, context=context@entry=0xffffffff81c80270, mask_=...,
local=local@entry=0xffffffff81103b10 <x86_invalidate_tlb(void*)>, context2=context2@entry=0xffffffff81c80270, flags=flags@entry=0) at kernel/smp.cpp:245
0xffffffff81103ebf in mmu_invalidate_range (addr=<optimized out>, pages=<optimized out>, mm=<optimized out>) at arch/x86_64/mmu.cpp:740
0xffffffff81076b01 in vm_invalidate_range (addr=<optimized out>, pages=<optimized out>) at kernel/mm/vm.c:2339
0xffffffff8107fe0b in tlbi_end_batch (tlbi=tlbi@entry=0xffffffff81c80308) at kernel/mm/memory.c:441
0xffffffff81080f58 in vm_mmu_unmap (mm=<optimized out>, addr=<optimized out>, pages=<optimized out>, vma=<optimized out>) at kernel/mm/memory.c:927
0xffffffff81078070 in __vfree (ptr=<optimized out>, is_mmiounmap=<optimized out>) at kernel/mm/vmalloc.cpp:335
0xffffffff81107757 in hpet_timer::~hpet_timer (this=0xffffd001270eef60) at arch/x86_64/hpet.cpp:108
unique_ptr<hpet_timer>::delete_mem (this=<synthetic pointer>) at include/onyx/memory.hpp:302
unique_ptr<hpet_timer>::~unique_ptr (this=<synthetic pointer>) at include/onyx/memory.hpp:377
hpet_init () at arch/x86_64/hpet.cpp:185
0xffffffff81030abc in do_init_level (level=3) at kernel/init.cpp:192
kernel_main () at kernel/init.cpp:233
0xffffffff810fdcd9 in x86_start () at arch/x86_64/boot.S:434
0x0000000000000000 in ?? ()

In other words, somewhere in the (fully no-op) HPET driver, the timer is
destroyed, the mmio mapping is freed, and the TLB shootdown hangs.

Further inspection reveals that the smp::sync_call_with_local is hanging
waiting for a CPU that never receives the message. Anecdotally looking
at the debug cpumask (with DEBUG_SMP_SYNC_CALL) shows it's usually the
last CPU to be brought up (e.g 0b1000 for a 4-cpu system).

The cause is simple: the BSP sets CPUs as online as soon as they set
boot_done. boot_done is set very early in boot. So, the following can
happen:

CPU 0              |   CPU N
wake up CPU N      |
                   |  runs early bootstrap
receive boot_done  |
set_online(N)      |
...                |
rest of init       |  smpboot_main()
...                |
smp_sync_call      |
[sends messages]   |
                   | lapic_init_per_cpu()

where CPU N misses the message sent from CPU 0.

Instead of being subject to this, spin in the BSP on the online mask,
and make the secondary core set itself online after doing bringup (just
before entering idle). This fully sidesteps the issue.

The fix only applies to x86 and should probably be done on the riscv
side as well.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Instead of conditionally creating it, create it for every inode. This
makes for less broken edgecases.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
writeback_inode() doesn't need ->writepages() if data isn't dirty.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
This fixes funny edgecases (which I no longer recall, because I've
carried this fix for a good while).

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Otherwise the destructor in exec_state() will free it.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Add group support and surrounding support code for rtnetlink
notifications.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Add infrastructure for broadcasting link status changes via workqueues
and netlink.

This is required by userspace daemons and tooling.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Add libmnl (per d24c88693be13ae77843e18e3388ca9142efdc31) and lightly
modified for Onyx.

All credit goes to Pablo Neira Ayuso & other authors.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Listen for, and track, link status changes and new devices via netlink
RTM_NEWLINK messages. This introduces a new dependency on libnl, a
lightweight library for manipulating netlink.

The inner workings are the following: instead of blindly using eth0
(despite it being present or not), create an rtnetlink socket bound to
the LINK group. netctld then initiates a GETLINK dump, while it listens
for changes on the LINK group.

If an interface was down, but now it's up, it starts configuration (via
normal dhcp, ipv6 procedures (already pre-existing). Not much changed
there. If an interface is new and up, it also starts configuration.

Doing interface teardown on RTM_DELLINK (or link down) is an exercise
left to the reader.

While at it, relicense the whole code to GPLv2.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Add peak RSS tracking, as per expected in getrusage, wait4, etc.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
Disable it.

Signed-off-by: Pedro Falcato <pedro.falcato@gmail.com>
@heatd
heatd enabled auto-merge (rebase) August 22, 2026 19:53
@heatd
heatd merged commit 2460502 into master Aug 22, 2026
17 checks passed
@heatd
heatd deleted the next branch August 22, 2026 22:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant