Imported from openpmix/prrte (
src/mca/odls/AGENTS.md). Install upstream withnpx skills add openpmix/prrte --skill odls. Copyright stays with the author.
AGENTS.md — The odls Framework (Daemon Local Launch Subsystem)
Orientation for AI agents and human contributors working in
src/mca/odls/. This is a map, not the rulebook: the authoritative
project guidance lives in the top-level AGENTS.md
and under docs/. When this file and those disagree,
the docs win — and please fix this file.
What this framework does
odls is the PRTE Daemon's Local Launch Subsystem. It is the code
that actually forks and execs the application processes on each node,
tracks them, signals them, reaps them via waitpid, and reports their
state transitions back into the job state machine. Where rmaps decides
which proc goes where (on the HNP), odls is what does the launch
(on every daemon).
Unusually for an MCA framework, odls has two very different jobs
depending on where the process runs:
| Role | Where | What odls does |
|---|---|---|
| HNP / DVM master | one process | Serializes the computed placement into a single launch message (get_add_procs_data) that is broadcast to all daemons. |
| prted (daemon) | every node | Parses that message (construct_child_list), works out which procs are local, and fork/execs them (launch_local_procs). |
The HNP is itself a daemon (vpid 0), so it also launches any local procs assigned to its own node — but its launch message construction is the part unique to it.
Place in the launch state machine
… → MAP → MAP_COMPLETE → SYSTEM_PREP → LAUNCH_DAEMONS → … → LAUNCH_APPS → RUNNING → …
▲
└── odls runs here
The flow that drives odls (see src/mca/plm/base/plm_base_launch_support.c
and src/prted/prted_comm.c):
- A job reaches
PRTE_JOB_STATE_LAUNCH_APPS. On the HNP,prte_plm_base_launch_apps()packs the daemon command (PRTE_DAEMON_ADD_LOCAL_PROCS, orPRTE_DAEMON_DVM_ADD_PROCSfor a fixed DVM) intojdata->launch_msg, then callsprte_odls.get_add_procs_data()to append the placement/regex/setup payload. - That call ends (asynchronously, after
PMIx_server_setup_applicationreturns) by activatingPRTE_JOB_STATE_SEND_LAUNCH_MSG, which xcastsjdata->launch_msgto every daemon over the RML. - Each daemon's
prted_comm.cdispatch seesPRTE_DAEMON_ADD_LOCAL_PROCSand callsprte_odls.launch_local_procs(buffer). launch_local_procs→construct_child_list(decode) →PRTE_ACTIVATE_LOCAL_LAUNCH→launch_local(per-app fork/exec) → each child transitions toPRTE_PROC_STATE_RUNNING.
The same module also fields the kill/signal daemon commands
(PRTE_DAEMON_KILL_LOCAL_PROCS, PRTE_DAEMON_SIGNAL_LOCAL_PROCS). Its
fifth entry point, restart_proc, has no caller anywhere in the tree -
nothing restarts a proc through the odls today, so that path is exercised
by nothing, including the swarm.
Directory layout
odls/
odls.h # module vtable (5 fn ptrs) + component typedef + version macro
odls_types.h # PRTE_DAEMON_* command flags; child-error pipe struct
base/
base.h # framework globals struct, base-fn prototypes, the two caddy classes,
# PRTE_ACTIVATE_LOCAL_LAUNCH / PRTE_ODLS_SET_ERROR macros
odls_base_frame.c # open/close/register; MCA params; class instances
odls_base_select.c # component selection (pick ONE, highest priority)
odls_base_default_fns.c # THE big one: build msg, parse msg, wireup, env setup, spawn, waitpid,
# kill, restart — everything a component reuses
odls_base_bind.c # prte_odls_base_set(): apply cpu/memory binding in the child pre-exec,
# proxy binding errors up the pipe
help-prte-odls-base.txt # xterm-related error text
pdefault/ # the only component (pri 10): real fork()/execve() launcher
Read odls.h and base/base.h first (the contract and the shared data
structures), then base/odls_base_default_fns.c, which is where almost
all real work lives. The pdefault component is a thin shell around the
base helpers — read it last.
The module contract
Every odls component fills in a prte_odls_base_module_t (declared in
odls.h) with five function pointers:
typedef struct prte_odls_base_module_1_3_0_t {
prte_odls_base_module_get_add_procs_data_fn_t get_add_procs_data;
prte_odls_base_module_launch_local_processes_fn_t launch_local_procs;
prte_odls_base_module_kill_local_processes_fn_t kill_local_procs;
prte_odls_base_module_signal_local_process_fn_t signal_local_procs;
prte_odls_base_module_restart_proc_fn_t restart_proc;
} prte_odls_base_module_t;
| Function | Signature | Runs on | Meaning |
|---|---|---|---|
get_add_procs_data |
(pmix_data_buffer_t *data, pmix_nspace_t job) |
HNP | Serialize the whole job (proc→node map, regex nodemap/procmap, personality, uid/gid, app-setup info) into data for broadcast. Returns PRTE_SUCCESS/error. |
launch_local_procs |
(pmix_data_buffer_t *data) |
daemon | Decode the message, build this node's child list, fork/exec the local procs. Returns PRTE_SUCCESS/error. |
kill_local_procs |
(pmix_pointer_array_t *procs) |
daemon | Kill the listed procs (NULL ⇒ all local procs). Escalates SIGCONT→SIGTERM→SIGKILL. |
signal_local_procs |
(const pmix_proc_t *proc, int32_t signal) |
daemon | Deliver signal to one proc (NULL ⇒ all local procs). |
restart_proc |
(prte_proc_t *child) |
daemon | Re-fork a single already-known child (fault recovery / comm-spawn restart). |
The return protocol is the ordinary PRRTE one: PRTE_SUCCESS or a
PRTE_ERR_*. Unlike rmaps, there is no "take next option" — one
component wins and owns every call. Errors on the daemon side almost
always end by activating a proc or job error state rather than
returning up the stack, because the launch runs asynchronously on the
event loop (PRTE_ACTIVATE_PROC_STATE(..., PRTE_PROC_STATE_FAILED_TO_LAUNCH),
PRTE_ACTIVATE_JOB_STATE(..., PRTE_JOB_STATE_NEVER_LAUNCHED)).
The version macro is PRTE_MCA_BASE_VERSION(odls).
Component selection is "pick one"
prte_odls_base_select() (odls_base_select.c) is the standard MCA
"select the single best component" pattern: it calls pmix_mca_base_select,
copies the winning module into the global prte_odls, and everything in
the tree calls through prte_odls.<fn>(). There is currently exactly one
component — pdefault, priority 10 — deliberately low so a
site-specific launcher could override it. The framework's open/select
logic only runs the launch machinery inside a daemon; a tool never
selects an odls module for launching.
What base/ provides — the heart of the framework
Because there is only one component and it delegates almost everything,
the base is the framework. A component supplies just the primitive
fork_local_proc (and the raw kill/signal syscalls); the base does
message construction, parsing, wireup, environment assembly, threading,
waitpid interpretation, and cleanup. Walk these in order.
1. Framework globals (odls_base_frame.c)
prte_odls_globals (prte_odls_globals_t in base.h) holds:
- No xterm state.
--xtermis a directive of a job, and a persistent DVM's daemons open this framework long before any job asks for it, so nothing about it can live here. It arrives as thePRTE_JOB_XTERMjob attribute (set bypmix_server_dyn.cfrom thePRTE_XTERM_RANKSspawn keyprun_common.cadds), andxterm_select()settles, per child and on the progress thread, whether that child runs underxterm -T "Rank N" [-hold] -e. It was once a DVM-wideprte_xtermglobal read here at open; when its MCA param was removed nothing replaced it, and--xtermparsed and did nothing. signal_direct_children_only— MCA flag controlling whether signals go to the child only or its whole process group.exec_agent— an optional wrapper command to exec instead of the app.
Forking is parallelized on the process-wide worker pool, which odls does
not own. The launch walk asks prte_worker_pool_assign()
(src/runtime/prte_worker_pool.h) for a
base per child and posts the fork there; the pool is sized by
prte_num_worker_threads (default 8) and is shared with the OOB, which puts
peer sockets on the same threads. odls used to carry a pool of its own —
ev_bases/ev_threads/num_threads/max_threads/cutoff/next_base, sized
from the first job it was handed — and every one of those fields and
parameters is gone. Two of its failure modes are worth remembering, because
anything that reintroduces per-subsystem sizing invites them back: the "have
we built a pool yet?" guard has to key off the pool itself and not off the
names of the threads (which stay NULL whenever the answer was "no dedicated
threads"), and a variable that is both the request and the answer must not
have its "you decide" sentinel restored by a teardown path, or a call arriving
during finalize builds a whole new set of real threads on the way out.
MCA params, all under prte odls base: signal_direct_children_only,
exec_agent, scatter_cpusets, and fork_publish_delay — a
fault-injection hook, like prte_daemon_fail, that stalls between the
fork and the store of the child's pid so the launch/reap race described
below can be reproduced deterministically from a single prterun. It is
not restricted to a debug build, on purpose: an optimized build is a
different race, so a hook that exists only in a debug one cannot say
anything about the build that ships. Framework open also unblocks SIGCHLD
(odls must see child deaths). Framework
close releases the global prte_local_children array.
This file also defines the two caddy classes (below) via
PMIX_CLASS_INSTANCE.
2. Building the launch message — prte_odls_base_default_get_add_procs_data()
Runs on the HNP only. Packs into the supplied buffer, in a strict order that the parser below must mirror (there is a literal comment in the source warning about this):
- The job being launched (
prte_job_pack). - A nodemap regex and a procmap (ppn) regex, generated by asking
the PMIx server (
PMIx_generate_regex/PMIx_generate_ppn) to compress the node names and the per-node rank lists. - Job info: personality, a per-job network allocation request, the
launching user's
uid/gid, and — if envars have not yet been harvested — aPMIX_SETUP_APP_ENVARSdirective.
Note what is deliberately not here. This message used to lead with an
int8 flag and, when a launch had brought new daemons in, a nested buffer
holding every other active job — so a freshly added daemon could resolve
their namespaces. That tied the size of the launch message to the number of
jobs resident in the DVM, sent them to every daemon rather than the ones that
needed them, and still did nothing for a daemon added by a bare elastic grow,
which launches no job and therefore sends no launch message at all. Those
jobs now ride with the nidmap at VM_READY
(prte_util_pack_job_catchup() in src/util/nidmap.c),
which is sent precisely when the daemon set changes.
And note what does the placement now. The job no longer packs a record
per process saying where it went: it packs a node map and one proc map per
app, and the receiver rebuilds each proc's rank, hosting daemon, app, app
rank and local rank from those — see
src/runtime/data_type_support/AGENTS.md.
What is left per proc is only what the maps cannot say (node rank, cpuset,
state, attributes), which is about 13 bytes a proc against the ~46 this
message cost per proc before any of this. plm_base_verbose 2 prints the
size; --rtos donotlaunch will size a job of any shape without launching
it.
And the cpuset does not go in this message at all — see 2a below.
It then calls PMIx_server_setup_application() asynchronously and does
not wait. It returns PRTE_SUCCESS immediately, having handed a
prte_odls_jcaddy_t to PMIx; the completion callback setup_cbfunc()
runs on the PMIx thread, so it only serializes the returned info into a
byte object and thread-shifts. _setup_complete(), back on
prte_event_base, packs that byte object onto jdata->launch_msg and
activates PRTE_JOB_STATE_SEND_LAUNCH_MSG — which is what actually
continues the launch. Nothing blocks: a PRTE_SUCCESS return here
means "the callback will fire", and an error return means it never will
(the caddy is released on the spot).
The status the callback is handed is a verdict, not a courtesy. PMIx
absorbs "not for me" (NOT_AVAILABLE, TAKE_NEXT_OPTION) below it, so a
non-success status means a network, GPU or programming-model component
really failed to set the job up. _setup_complete() fails the job
(NEVER_LAUNCHED) rather than launching it without whatever that component
was to provide; _local_support_complete() does the same on a daemon for
PMIx_server_setup_local_support().
2a. The half addressed to one daemon — prte_odls_base_send_cpuset_slices()
A proc's binding is read by exactly one daemon, the one that forks it, and
it is the largest of the three things a proc still costs the launch message
(~6 B/proc raw on a 176-core node, ~40% of the message). Broadcasting it
put every daemon's bindings on every link of the tree. So the launch message
is packed PRTE_JOB_PACK_NO_CPUSETS and each daemon is sent its own
bindings point to point, from prte_plm_base_send_launch_msg() immediately
before the broadcast. No new data movement was needed: prte_rml_get_route()
sends each one toward its target, so the bytes leaving the master are the
sum of the slices rather than the sum replicated per daemon.
The slice is nspace, a count, and then (rank, cpuset) per proc. The
rank is carried, not implied by position — the receiver would otherwise
have to reproduce the order the sender's loop happened to walk in, and
getting that wrong binds processes to each other's cpus without failing
anywhere.
The two halves are independent messages and arrive in either order.
Neither order is the rare one — the broadcast takes more hops, the slices
are sent one at a time — so both are handled the same way, by a rendezvous
on prte_odls_globals.pending_slices: whichever arrives first parks, and
whichever arrives second finds it and completes the launch.
construct_child_list therefore does not always register the nspace
before it returns; when it parks, prte_odls_base_recv_cpuset_slice() is
what calls start_registration() later. Both run on prte_event_base, so
the list needs no lock.
Three cases that must not wait, and each is a hang if you get it wrong:
- the master, which keeps its own fully populated job object and is never sent a slice;
- a daemon hosting none of the job's procs — the broadcast reaches every daemon, the slices only the ones in the map. It has nothing to bind, so it waits for nothing (and discards anything parked for that job);
PRTE_JOB_PACK_ALL, the shape a spawn request and the job catchup use, where the cpusets were in the buffer all along. The receiver keys off the mode byte it unpacked, not off its own MCA parameter, so a daemon can never wait for a slice the master did not send.
A parked half belongs to its job, and goes with it. The caddy carrying
a parked launch holds its own reference on the job, and the list entry owns
the caddy, so dropping the entry releases both.
prte_odls_base_discard_slices() does that, and prted_comm.c calls it on
every daemon when the DVM cleans the job up — before looking the job up,
because a slice whose launch message never came has no job to find. The
case it closes is a slice that is late rather than lost: a relay on its
route dies, the job is aborted and cleaned up, and the RML replays the slice
afterwards. A launch still parked for it would otherwise hold a job the
cleanup freed, and the replayed slice would register and fork the procs of
a job that is over. The lookup is
strict (PMIX_CHECK_NSPACE_STRICT): an undecodable namespace must not
match, and so discard, somebody else's parked launch.
--prtemca odls_base_scatter_cpusets 0 turns it off and broadcasts the
bindings as before; it is an A/B switch, not a supported difference in
behavior.
What this costs: a daemon no longer holds the binding of a proc it
does not host, so a PMIx_Get for a remote proc's PMIX_CPUSET or
PMIX_LOCALITY_STRING becomes a direct modex to the daemon that does -
which answers it out of its own job object without waiting on the process.
The value is unchanged; what it costs is one round trip, once, since PMIx
caches the reply. Nothing else on a proc is affected, and a proc on the
asker's own node - the only place a locality string means anything - is
held locally as before.
Three states that had to be told apart to make that referral safe, and
the whole of src/prted/pmix/AGENTS.md's
dmodex section is about them. The hosting daemon may not have the answer
yet (its own slice is still in flight - prte_odls_base_awaiting_cpusets()
says so, and the request waits); it may never have it again (its share of
the job finished and it released the job object - it answers NOT_FOUND);
or it may simply not have been reached by the launch message yet (it waits).
Confusing the second with the third is a hang with nothing logged.
3. Parsing the message and wiring up — prte_odls_base_default_construct_child_list()
Runs on every daemon (including the HNP). This is the mirror image of
get_add_procs_data, and the single most important function to understand:
- Unpacks the job to launch. On the master it throws away the unpacked copy
and fetches the fully-populated local
prte_job_t; on a daemon it keeps the unpacked copy, creates amapif needed, and resolves the job's schizo personality viaprte_schizo_base_detect_proxy(). - Unpacks the optional app-setup byte object, and folds any
PMIX_SET/ADD/UNSET/PREPEND/APPEND_ENVARitems into the job attributes (prepended, so they apply before launch). - Wireup loop: for every proc in the job, connect it to its node via
the parent daemon (
daemons->procs[pptr->parent]->node), add the node to the job map once (guarded byPRTE_NODE_FLAG_MAPPED, which is then reset), and — crucially — decide locality: ifpptr->parent == PRTE_PROC_MY_NAME->rank, the proc is mine. Local procs are retained onto the globalprte_local_childrenarray, flaggedPRTE_PROC_FLAG_LOCAL, counted intojdata->num_local_procs, and their app is flaggedPRTE_APP_FLAG_USED_ON_NODE. - Registers the nspace with the PMIx server
(
prte_pmix_server_register_nspace) and returns. The registration completes asynchronously; aprte_odls_jcaddy_tcarries the launch across it, andjob_reg_join()is the join point. It runsPMIx_server_setup_local_supportif setup info was present, starts the spawn threads, and firesPRTE_ACTIVATE_LOCAL_LAUNCHonce everything has reported. Nothing blocks here either.- The join is guarded by a sentinel:
cd->pendingstarts at 1 and is only decremented by thejob_reg_join()call at the bottom of the function, so the count cannot reach zero — and release the caddy — part-way through issuing the registrations. Keep that shape if you add another registration. - The registration callback (
_job_reg_complete) is invoked on the PRRTE progress thread:prte_pmix_server_register_nspacethread-shifts before calling back. That is what makes the unlockedpendingbookkeeping safe.
- The join is guarded by a sentinel:
- On any failure it calls
fail_local_procs()and then activatesPRTE_JOB_STATE_NEVER_LAUNCHEDso the HNP doesn't hang waiting for a daemon that silently died. The activation alone is not enough, which is why the two are always paired: what the prted errmgr sends the HNP is the state of this daemon's local children of that job, so children left in whatever state they were unpacked in read at the HNP as "nothing wrong here" — andfailed_start()only retires procs it finds already markedPRTE_PROC_STATE_FAILED_TO_START, so without the marking they sit inprte_local_childrenforever and this daemon's own termination accounting waits on procs that do not exist.fail_local_procs()marks them, setsPRTE_PROC_FLAG_IOF_COMPLETEandPRTE_PROC_FLAG_WAITPID(nothing was forked and no stdio was ever opened, so both are complete by definition), adds any that never reachedprte_local_children, and resetsnum_local_procsto match. It is a no-op on the master, which reports to nobody and holds the authoritative job object. - The namespace must be free. The launch message is what gives a
daemon a job, so
prte_set_job_data_object()has to succeed; if a copy is already on file, the procs just wired up belong to an object nobody will find by name —launch_locallooks the job up, gets the other copy, sees no local procs, and forks nothing. A grow's catch-up used to plant exactly such a copy for every job the elastic launch fence was holding; seePRTE_JOB_FLAG_LAUNCH_PENDINGinsrc/util/AGENTS.md. A refusal is reported like any other failure here. jdataat theREPORT_ERRORlabel must be the job we were told to launch, or NULL. The prior-jobs loop that used to sit at the top of this function decoded other jobs, and reusedjdatato do it — so a failure later in the loop drove the wrong (possibly already-released) object toNEVER_LAUNCHEDand the job actually being launched was never reported at all. That loop has moved to the nidmap, where it decodes into its own variable for the same reason; keep any new decoding here to its own.
4. Kicking off the fork — PRTE_ACTIVATE_LOCAL_LAUNCH and prte_odls_base_default_launch_local()
The component's launch_local_procs finishes by invoking the
PRTE_ACTIVATE_LOCAL_LAUNCH(job, fork_local_proc) macro (in base.h),
which allocates a prte_odls_launch_local_t caddy, stashes the component's
fork_local primitive on it, and posts prte_odls_base_default_launch_local
to prte_event_base.
prte_odls_base_default_launch_local() is the per-node launch driver:
-
Records a baseline
getcwd(it willchdiraround per app and must return here). -
Enforces the system limits on total children and open file descriptors; if over budget it retries via a
PRTE_DETECT_TIMEOUTtimer (up to a few times) rather than failing outright. -
For each app used on this node: sets up the working directory (
setup_path, honoringPRTE_APP_SSNDIR_CWD/PRTE_APP_USER_CWD), mergesprte_launch_environintoapp->env, applies env directives (process_envars— the SET/ADD/UNSET/PREPEND/APPEND handling, with app attributes trumping job attributes), calls the schizo'ssetup_fork, places prepositioned files in that working directory (prte_filem, which reads theapp->cwdsetup_pathjust resolved — so the order of those two calls is load-bearing), checks the executable, and applies resource limits.process_envarsowns the envar directives — all of them. It runs for every personality whatever that personality'ssetup_forkdoes, which is why the schizo hook deliberately does not repeat them: applyingPREPEND/APPENDtwice is not a no-op (it duplicates every entry a user prepends ontoPATH), andprte_schizo_base_setup_forkused to do exactly that. The schizo hook is left with the PMIx prefix only.Order is the contract. The list is walked front to back, and that order is the order the user's directives take effect in:
SETreplaces a value outright whilePREPEND/APPENDedit the one already there, so--prepend-env FOO[:] x --set-env FOO=1leavesFOO=1and the reverse order leavesFOO=x:1. Neither is a merge policy we get to choose — the user said what to do and in what sequence. What builds the list in that sequence isprte_append_attribute()inpmix_server_dyn.c; theprte_prepend_attribute()calls inconstruct_child_listare the deliberate exception, putting the valuesPMIx_server_setup_applicationreturned in FRONT so the user's own directives, which arrive on the job already, are applied afterwards and win. That block is walked in reverse for the same reason: prepending a block one entry at a time reverses it.Two shapes to keep in mind when editing it:
UNSETis aPMIX_STRING, not apmix_envar_t— it names a variable and carries no value. It has to be handled before thePMIX_ENVAR != attr->data.typefilter that guards the rest of the loop, or--unset-envsilently does nothing. A trailing*makes the name a prefix.- Match an environment entry's name up to and including the
=. A barestrncmpof the name's length matches any variable that merely starts with it, so prepending ontoPATHwould editPATHEXTinstead — whichever the environment happens to list first. That is whatenvar_value()is for.
-
For each local child of that app in
INIT/RESTARTstate: registers thewaitpidcallback (prte_wait_cb→prte_odls_base_default_wait_local_proc), setsPRTE_PROC_FLAG_ALIVE, allocates aprte_odls_spawn_caddy_t, sets up IOF (prte_iof_base_setup_prefork/setup_parent), picks the next event base from the thread pool, and postsprte_odls_base_spawn_procto it.
STOP_ON_EXEC caveat: if PRTE_JOB_STOP_ON_EXEC is set (debugger
attach), the fork is forced onto prte_event_base rather than a worker
thread, because the ptrace tracer must be the same thread that later
detaches — see the long comment near the thread-selection code.
5. The spawn step — prte_odls_base_spawn_proc()
Runs on the chosen event base. This is the last common code before the component's raw fork:
- Honors
PRTE_JOB_DO_NOT_SPAWN(mapping-only "donotlaunch" jobs): just mark the childTERMINATEDand return. - Calls
PMIx_server_setup_fork()to inject the PMIx client environment. - Resolves the actual command/argv: normal app, or the
--xtermwrapperxterm_select()put on the caddy (cd->xterm_argv), or a per-jobPRTE_JOB_EXEC_AGENT, or the globalexec_agent; optionally index-suffixesargv[0]with the rank (PRTE_JOB_INDEX_ARGV). - Calls the component's
cd->fork_local(cd)— the actualfork/execve. - On success stores the pid in PMIx (on the master), and in every case
hands its verdict to
spawn_doneonprte_event_base, which records it on the child and activatesPRTE_PROC_STATE_RUNNINGor the failure state. See "The fork does not write the child" below.
6. Applying binding — prte_odls_base_prepare_binding() + prte_odls_base_set() (odls_base_bind.c)
Binding is split across the fork so the child stays async-signal-safe:
prte_odls_base_prepare_binding(cd)runs in the parent, inspawn_procjust beforefork_local. It does everything that allocates, parses, or prints: it parses the proc's computedchild->cpuset(the hwloc bitmap string the mapper produced) into a storedhwloc_cpuset_t, classifies the binding, precomputes the memory-binding policy, emits--report-bindingsoutput and the "incorrectly bound" warning, and — where the platform hassched_setaffinity(PRTE_HAVE_SCHED_SETAFFINITY) — precomputes a rawcpu_set_taffinity mask. All of this is stashed on the caddy.prte_odls_base_set(cd, write_fd)runs in the forked child, in the async-signal-safe window beforeexecve. It only issues the bind syscalls: a baresched_setaffinitywith the precomputed mask on Linux, orhwloc_set_cpubindas the#elsefallback (macOS and other platforms withoutsched_setaffinity), and a bareset_mempolicywith the precomputed mode and nodemask (hwloc_set_membindas the#else). It allocates nothing and renders nothing.
Because the child is not a real PRTE process — and runs in that
async-signal-safe window — it cannot use normal error reporting or
render a show_help message (that allocates, reads the help file, and
scans directories, any of which can deadlock in a forked child). It
reports a fixed-size code-plus-errno record up the pipe via
prte_odls_base_child_fail (fatal, _exits) / prte_odls_base_child_warn
(non-fatal, returns) — the prte_odls_pipe_err_msg_t /
prte_odls_child_err_t types in odls_types.h — and the parent renders
the human-readable diagnostic. Whether a binding failure is fatal or a
warning depends on PRTE_BINDING_REQUIRED and PRTE_BINDING_POLICY_IS_SET
(a required, explicitly-requested binding that fails kills the child; a
defaulted one degrades to a warning). If the proc has no cpuset but the
daemon itself is bound, the proc is "freed" to all allowed cpus.
The child calls into hwloc on no platform that has the syscalls. Both
bind calls hwloc offers allocate — hwloc_set_membind several times, on its
way to the single set_mempolicy(2) it issues underneath — and a malloc
in a forked child deadlocks whenever another thread of the daemon held the
arena lock at fork time. The daemon is multi-threaded (the PMIx progress
thread, plus the shared worker pool the forks themselves run on), so that
window is real. prte_odls_base_prepare_mempolicy() therefore does hwloc's
work in the parent: the cpuset→nodeset conversion (hwloc_set_membind →
hwloc_fix_membind_cpuset → hwloc_fix_membind), the policy→MPOL_*
translation, and the kernel nodemask
(hwloc_linux_membind_mask_from_nodeset). It is compiled on every platform,
not only the ones that can issue the syscall, so that the unit test covers
it wherever it runs; PRTE_HAVE_SET_MEMPOLICY (config/prte_check_mempolicy.m4)
gates only the syscall, with the old hwloc_set_membind left as the #else.
Three decisions the parent makes so the child has only the one call left:
- A topology that is not this machine gets no memory binding at all.
Under
hwloc_use_topo_filethe topology describes some other machine, and hwloc gives such a topology no-op binding hooks that report success without touching anything. Aiming a realset_mempolicyat another machine's NUMA numbering would not be reproducing that — it would be binding a process by numbers that mean nothing here — sodo_membindis cleared. (Note the asymmetry with cpu binding, which goes throughsched_setaffinityand is applied from such a topology.) - A platform with no memory binding is answered as hwloc would, with
ENOSYSrecorded incd->membind_prep_errnofor the child to report. The support bits are read throughhwloc_topology_get_support()and are only meaningful for a this-system topology —set_topology()insrc/hwloc/hwloc_base_util.casserts them back on for an imported one, so they say nothing there. - Only
MPOL_DEFAULTandMPOL_BINDare expressed, becauseprte_hwloc_base_maponly ever asks forHWLOC_MEMBIND_DEFAULTorHWLOC_MEMBIND_BIND|STRICT. Anything else recordsENOSYSrather than guessing. If a third memory policy is ever added, this is what has to grow with it.
A failure to apply the default policy is not a memory-binding failure.
With mem_alloc_policy=none — the default — the call is a reset to the
system default, not a binding: nothing ends up bound to the wrong place and
nothing is degraded, so it is not reported at all, whatever the errno and
whatever mem_bind_failure_action says (that parameter governs an explicit
binding, and its own help text says so). The rule used to be narrower —
ENOSYS alone — which covered the platform that cannot bind memory but
not the one that refuses to: Docker's default seccomp profile denies
set_mempolicy with EPERM, and every bound launch inside a container
announced that "the memory was left unbound", aborting the job outright
under mem_bind_failure_action=error.
7. Reaping children — prte_odls_base_default_wait_local_proc()
The waitpid callback, registered per child and fired by
src/runtime/prte_wait.c when SIGCHLD is reaped. It decodes
proc->exit_code (the raw wait status) into a proc state:
WIFEXITED+ zero ⇒PRTE_PROC_STATE_WAITPID_FIRED.WIFEXITED+ nonzero, withPRTE_JOB_ERROR_NONZERO_EXITset ⇒PRTE_PROC_STATE_TERM_NON_ZERO.- Exited "normally" but never did the required PMIx init/finalize sync ⇒
PRTE_PROC_STATE_TERM_WO_SYNC(checked againstPRTE_PROC_FLAG_REG/PRTE_PROC_FLAG_HAS_DEREGandprte_allowed_exit_without_sync). WIFSIGNALED⇒PRTE_PROC_STATE_ABORTED_BY_SIG, and the exit code is rewritten tosigno + 128(shell convention, soprogandprun progagree).- Proc that called
prte_abort⇒PRTE_PROC_STATE_CALLED_ABORT; a proc ordered dead (KILLED_BY_CMD) is passed straight through. - STOP_ON_EXEC (
WIFSTOPPED+SIGTRAPunderPRTE_JOB_STOP_ON_EXEC): this is the debugger-attach stop. Detach with SIGSTOP so the child stays parked for the debugger, re-register the waitpid, firePRTE_PROC_STATE_READY_FOR_DEBUG, and do not fall through to exit handling. This detach must run onprte_event_base(see the fork thread-affinity note above).
It ends at MOVEON: by cancelling the wait tracker and activating the
computed proc state.
8. Kill / signal / restart
-
prte_odls_base_default_kill_local_procs()— walks the requested procs againstprte_local_children, closes stdin IOF, cancels the waitpid (to avoid races), then escalates SIGCONT → SIGTERM → SIGKILL withnanosleepgaps, marking eachKILLED_BY_CMD. It calls the component's rawkill_local(pid, signum).A child can be ordered to die mid-fork. It is
ALIVEwith no pid from the momentlaunch_localdispatches it until its fork reports back, and a kill in that window - an aborted job, one of whose ranks failed while its siblings on the same node were still being forked - has nothing to signal. It must not be treated as never having started either: the fork goes ahead regardless, and a child recorded as terminated then ran to its own end while the aborted job waited on it. Such a child is flaggedPRTE_PROC_FLAG_KILL_PENDING, andspawn_donedelivers the kill once the pid exists, batching the children that report while one batch is pending (the escalation sleeps on the progress thread, so one call per child would stall the daemon per child).--prtemca odls_base_fork_publish_delay 3000000withprte_num_worker_threads 1and a two-rank job whose rank 0 exits non-zero holds rank 1 in exactly that window. -
prte_odls_base_default_signal_local_procs()— finds the target child (or all) and calls the component's rawsignal_local(pid, signum). -
prte_odls_base_default_restart_proc()— resets a single known child's state/flags and re-dispatches it throughprte_odls_base_spawn_proc(same caddy/thread/IOF machinery as a first launch).
Key data structures
| Type | Where | Purpose |
|---|---|---|
prte_odls_base_module_t |
odls.h |
The 5-pointer vtable; the selected one lives in the global prte_odls. |
prte_odls_globals_t / prte_odls_globals |
base.h / frame.c |
Framework-wide state: exec agent, signal policy, the cpuset-slice rendezvous list, the fork-publish fault injection. |
prte_local_children |
src/runtime/prte_globals (a pmix_pointer_array_t) |
The daemon's authoritative list of the procs it launched — every base fn iterates it. Allocated at framework open, released at close. |
prte_odls_spawn_caddy_t |
base.h |
Per-child fork caddy: cmd, wdir, argv, env, jdata, app, child, IOF opts, and the fork_local fn ptr. Carries ev for thread-shifting. Heap-allocated, released after spawn. |
prte_odls_launch_local_t |
base.h |
Per-node "start launching job J" caddy carried by PRTE_ACTIVATE_LOCAL_LAUNCH; holds job, fork_local, and a retries counter for the sys-limit backoff. |
prte_odls_pipe_err_msg_t / prte_odls_child_err_t |
odls_types.h |
Fixed-size record written up the child→parent pipe (fatal flag + exit status + failure code + errno). Carries no strings and needs no allocation, so it is safe to emit from the async-signal-safe window before execve; the parent renders the show_help diagnostic from the code and errno. |
PRTE_DAEMON_* command flags |
odls_types.h |
The daemon command byte that leads every RML control message to a prted (ADD_LOCAL_PROCS, KILL, SIGNAL, EXIT, …). |
Threading model
- Message build/parse, wireup, waitpid interpretation, kill/signal, and
state activation all run on the progress thread (
prte_event_base) — the normal PRRTE event-driven model. - Only the fork/exec spawn step may be off-loaded to the process-wide
worker pool (
prte_worker_pool_assign()) to parallelize launching many procs. Each spawn is a self-contained caddy handed to one worker base. The pool is shared with the OOB — a worker base servicing your fork may also be servicing a peer socket. - The fork does not write the child. "Its own child" is not the same
as "a child nobody else touches": the SIGCHLD reaper, the IOF read
handlers and the daemon command processor all reach that
prte_proc_tfrom the progress thread while the fork is in flight. A state written on the worker can overwrite a termination the progress thread has just recorded, andPRTE_FLAG_SET/UNSETare read-modify-write on a shared bitmask, so two threads touching different bits lose one. So everything the child needs cleared before a launch is cleared bylaunch_reset()on the progress thread, everything the launch learns is carried back byspawn_report()and recorded byspawn_done, and the only field the worker writes is the pid - published behind the component's gate for the reaper. The same goes for the job's attribute list, which the progress thread can append to mid-launch (the prted errmgr'sFAIL_NOTIFIED):spawn_caddy_resolve()reads what the fork path needs onto the caddy before dispatch, and nothing past the dispatch callsprte_get_attributeon the job. SIGCHLDmust stay unblocked (done at framework open); child death is delivered throughsrc/runtime/prte_wait.c, which fires the registeredwait_local_proccallback onprte_event_base.- No base function blocks.
get_add_procs_dataandconstruct_child_listboth hand their work to PMIx and return; the launch is carried across the gap by aprte_odls_jcaddy_t— which holds its own reference on the job, since the gap can outlast it — and resumed by a thread-shifted completion handler. Neither uses aprte_pmix_lock_t, and neither should grow one — they run onprte_event_base, which is the only thread that can run the handler they would be waiting for (see the top-levelAGENTS.md, "thread-shift every PMIx callback").
Gotchas when editing
- Pack/parse symmetry is sacred.
get_add_procs_dataandconstruct_child_listare a hand-matched serializer/parser pair. Any change to the packed order/type in one must be mirrored in the other, or daemons will mis-decode the launch message and the job hangs or crashes. The source says so in capitals — heed it. - Locality is
parent == my vpid. A proc is "local" iff itsparentdaemon vpid equals this daemon's rank. Getting the wireup wrong silently launches procs on the wrong node or not at all. - Raw back-pointers on unpacked objects are NULL on the daemon. Fields
like
prte_app_context_t.jobare only set on the HNP (iness/hnpand the dynamic-spawn path); they are not serialized into the launch message, so on any daemon that rebuilt the job viaconstruct_child_listthey are NULL. Never dereferenceapp->job(or similar back-pointers) in a daemon-side path —setup_pathdid, and--preload-binary(which setsPRTE_APP_SSNDIR_CWD) segfaulted on it. Get the job from the caller, which already holdsjobdat, or viaprte_get_job_data_object(nspace). prte_local_childrenis the single source of truth on a daemon. Adding/removing a child there, and itsPRTE_PROC_FLAG_*flags (LOCAL,ALIVE,WAITPID,IOF_COMPLETE,REG), gate the whole lifecycle. A child is only fully released once bothWAITPIDandIOF_COMPLETEare set.- This framework reports that the waitpid half is done; it does not set
the flag and it does not decide the proc is finished. Wherever the odls
knows a child will produce no further
SIGCHLD— the reaper, and both arms of the kill path — it activatesPRTE_PROC_STATE_WAITPID_FIRED, andprte_state_base_join()is the only thing that setsPRTE_PROC_FLAG_WAITPIDand the only thing that joins the two halves. A diagnosis is a separate activation, made first and only by whoever actually diagnosed the proc; the kill path in particular must not re-activate a diagnosis some other code made. Setting the flag by hand and re-activating the proc's existing state instead is what left an aborting rank unretired, its daemon's batchedUPDATE_PROC_STATEunsent, andprterunhung with every daemon alive. - Failure means activating a state, not returning. On the daemon side
the launch is asynchronous; report errors with
PRTE_ACTIVATE_PROC_STATE(FAILED_TO_LAUNCH/FAILED_TO_START)orPRTE_ACTIVATE_JOB_STATE(NEVER_LAUNCHED)so the HNP can react — don't just bubble anrcup into the event loop. - Use
prte_show_help(), notpmix_show_help(). Aprtedcannot deliver its own help text — PMIx'splog/stdfdonly writesstderrfor a client or tool, and for a server it hands the message to an IOF sink that does not exist — so everyshow_helpa daemon rendered used to be dropped on every node but the head one.prte_show_help()(src/util/prte_show_help.h) is a drop-in that relays to the HNP. Every call site in this framework is converted; keep new ones that way. - The child cannot log normally — or render
show_help. Betweenforkandexecveonly async-signal-safe calls are permitted, so the child must not allocate, use stdio, scan/proc/self/fd, or callshow_help(all of which can deadlock in a forked child). It reports a fixed-size code-plus-errno record up the pipe (prte_odls_base_child_fail/prte_odls_base_child_warn) and the parent renders the message; never call ordinary PRRTE logging there. - STOP_ON_EXEC pins the tracer thread. Both the fork (in
launch_local/restart_proc) and the ptrace detach (inwait_local_proc) must happen onprte_event_base; do not "optimize" them onto a worker thread. chdirbookkeeping.launch_local/restart_procbounce the daemon's cwd per app and must alwayschdirback tobasedirbefore returning.launch_localestablishesbasedirwithgetcwdat entry; if that fails it must not fall through to the sharedchdir(basedir)cleanup (the buffer is uninitialized) — bail directly.- Every fatal launch error aborts the whole job. All fatal paths in
launch_local—setup_path,setup_fork,filem,check_context_app,init_sys_limits, thechdir(basedir)-back failure, the sys-limit giveups, and the per-child IOF-setup failures — use the same idiom: flag the directly-affected procs withPRTE_ODLS_SET_ERROR(job, rc, j)(or activate the single failing child) and thenPRTE_ACTIVATE_JOB_STATE(jobdat, PRTE_JOB_STATE_FAILED_TO_LAUNCH)beforegoto GETOUT. Failing only this app's procs is not enough:goto GETOUTskips every later app, whose procs stay inINITforever, and on a real daemon the errmgr only reportsFAILED_TO_LAUNCHoncenum_terminated == num_local_procs— so the job hangs. The job-state activation is what drives the prompt, uniform teardown. Keep new fatal paths on this idiom. - The
childloop variable is only valid inside the per-child loop. Inlaunch_local,childisNULLat function entry and, in the per-app loop (before the inner per-child loop runs), is eitherNULL(first app) or a stale proc left over from the previous app. Error paths in the per-app section (e.g. thechdir(basedir)-back failure) must not dereferencechild— use the job-abort idiom above. - Every exit from
spawn_procmustPMIX_RELEASE(cd). The spawn caddy isPMIX_NEW'd inlaunch_localand owned byspawn_proc; the success tail, theerrorout:path, and the earlyPRTE_JOB_DO_NOT_SPAWNreturn all have to release it, or every donotlaunch/mapping-only proc leaks a caddy (plus itswdirstring). Each of them also owes exactly onespawn_report(), or the child's launch is never accounted for. setup_pathwrites through toapp->cwd.launch_localcalls it assetup_path(jobdat, app, &app->cwd), so it overwrites (and must first free) the app's existingcwd;restart_procinstead passes a local temp. Don't assume*wdirstarts NULL.- One cleanup label, not scattered
returns.get_add_procs_dataandconstruct_child_listbuildpmix_info_tarrays andPMIX_INFO_LISTs that every exit has to release. Prefer a singlegoto REPORT_ERRORover a barereturnin the middle — a strayreturnthere both leaks and, on the daemon side, skips theNEVER_LAUNCHEDactivation that keeps the HNP from hanging. - A child must not be able to die before its pid is recorded. The fork
happens on a worker thread, and the
SIGCHLDreaper on the progress thread has nothing butchild->pidwith which to attribute what it reaps — a status it cannot attribute is discarded, and the proc then never leavesRUNNINGand the job never completes.fork_local_procholds the child on a second pipe until the store is done; seepdefault/AGENTS.md, "The gate". A fork primitive added to a new component owes the same. - A launch that never forked owes a
prte_wait_cb_cancel.launch_localflags a childALIVEand registers its waitpid before the IOF setup and the dispatch tospawn_proc, so anything that fails in between leaves a wait tracker — holding a reference on the child — waiting for a pid that will never exist.spawn_donekeys the cancel off0 >= child->pidfor any launch that did not endRUNNING, which is exactly the no-process case (never forked, or the fork itself failed and stored -1); a post-fork failure keeps its registration, because that child does exist and itsSIGCHLDis still coming. - Never hand a pid of 0 to the kill/signal primitives. Both component
primitives turn a pid into a process group (
-pid) so a signal reaches whatever the app itself spawned — which makespid == 0catastrophic rather than merely useless:kill(0)andkill(-0)are both "every process in the caller's group", i.e. the daemon signals itself, its other children, and underprterunthe launching tool. A child sits at pid 0 for real windows — before its fork, after a failed launch, and oncekill_local_procshas cleared it — andprted_comm.casks the odls for every local child of the named job, by name, so those procs are handed to us as a matter of course. Both branches ofsignal_local_procsgate on0 >= child->pid && ALIVE, and both primitives refuse a non-positive pid on their own. Keep both layers: the cost of missing it once is a node's daemon killing itself. - A failing proc's
exit_codeis deliberately a union of the PMIx and PRRTE numbering schemes — do not "fix" it by converting. Everything that reachesPRTE_ODLS_SET_ERROR(or setschild->exit_codedirectly) ends up switched on byprte_render_launch_failure()insrc/runtime/prte_quit.c, and that switch has arms for both:PMIX_ERR_JOB_EXE_NOT_FOUND,PMIX_ERR_EXE_NOT_ACCESSIBLE,PMIX_ERR_JOB_WDIR_NOT_FOUND,PMIX_ERR_SYS_LIMITS_*sit next toPRTE_ERR_FAILED_GET_TERM_ATTRSand friends. So the statusespmix_util_check_context_app()andpmix_util_check_context_cwd()return must be stored as they are. Running them throughprte_pmix_convert_status()is exactly the sort of thing the boundary rule in the top-levelAGENTS.mdasks for, and here it is wrong: it collapses them onto a genericPRTE_ERRORand the user getsError code: -3001 / Error name: Errorinstead of being told which executable or directory was the problem. (Tried; the swarm'stest_odlscaught it.) ADDis notSET.PRTE_JOB_ADD_ENVAR/PRTE_APP_ADD_ENVARmean "set this envar unless it already has a value" — that is whatsrc/util/attr.hsays and what thePMIX_ADD_ENVARthey are translated from says.prte_odls_base_process_envarspassesoverwrite = falsefor them andtrueforSET; the two lines look identical otherwise, which is how they came to be the same.- Standard PRRTE rules apply:
prte_config.hfirst, braces on every block,NULL ==/constant-on-left comparisons,PRTE_ERROR_LOGfor unexpected errors, no new compiler warnings.
Debugging
prte --prtemca odls_base_verbose 5 ... # trace child-list build, dispatch, waitpid
prte --prtemca odls_base_verbose 10 ... # + sys-limit checks, per-thread dispatch
prte --prtemca odls_base_verbose 20 ... # >15 dumps the exact argv/env being exec'd
prte --prtemca state_base_verbose 5 ... # see LAUNCH_APPS / RUNNING transitions odls drives
prun --xterm 0,1 ... # route ranks 0,1 to xterm windows (see xterm_select)
prun --report-bindings ... # print each child's applied binding (odls_base_bind.c)
Useful tuning params: --prtemca prte_num_worker_threads N (the process-wide
worker pool the forks are dispatched to; 0 forks on the main progress
thread), --prtemca odls_base_exec_agent CMD (wrap every exec),
--prtemca odls_base_signal_direct_children_only 1 (don't signal the
child's whole process group).
Testing
The fork/exec/waitpid/kill lifecycle runs only against real child
processes inside a live DVM, so most of odls is covered by the
integration harness — e.g. prterun -n 4 hostname (basic launch),
MPMD (prterun -n 2 echo a : -n 2 hostname, which walks the per-app
setup_path loop), and a non-zero exit (prterun -n 2 sh -c 'exit 3',
which exercises wait_local_proc's status decode).
Multi-node — test_odls in the swarm harness
contrib/dockerswarm/run-tests.sh
carries a test_odls for the half that only exists once a remote
daemon has to decode the launch message and fork against its own copy:
- the envar directives (SET/ADD/UNSET/PREPEND/APPEND) end to end, via the
envspawnclient — they have no command-line surface at all, so aPMIx_Spawnis the only way to reach them, andenvspawnpins its child to a node its parent is not on so the fork happens elsewhere; - the exec agent, which each daemon reads from its own MCA state;
- a bad
execveon a remote node — the child cannot render its own message, so this is the proof that the parent daemon'srender_child_msgreaches the tool (including thestat-based "it is a directory" case); - an MPMD launch whose first app is bad, which is the regression guard for the job-abort idiom below: the observable is that it terminates rather than hanging;
- the waitpid decode arriving from a remote node (
exit 3stays 3, a SIGTERM death becomes 143); - the SIGCONT→SIGTERM→SIGKILL escalation, against a child that traps SIGTERM.
Without a DVM
What runs without one is the structural contract plus the pieces that
are pure functions, in
test/unit/odls/test_odls.c (wired into
make check): the pdefault module vtable is fully wired and reuses the
base get_add_procs_data; the component names itself "pdefault"; the
PRTE_DAEMON_* command bytes are pairwise unique; the
prte_odls_child_err_t fatal/warn split holds (NONE == 0, every warn
code sorts after every fatal code); and the two caddy classes
(prte_odls_spawn_caddy_t, prte_odls_launch_local_t) construct with
the documented NULL/zero defaults and destruct cleanly (both the
all-NULL and the fully-populated paths).
Two cases go further than structure, because the code under them is a pure function of its inputs:
prte_odls_base_process_envars— exported (rather than file-static) precisely so this test can drive it: SET vs ADD, the front-to-back ordering,UNSETwith and without a trailing*, the "match up to and including the=" rule (a directive onPATHmust not editPATHEXT), and app-trumps-job.- The child→parent pipe record —
child_warnacross a real pipe, andchild_failacross a realfork, checking both the decoded record and that the child died with the exit status it was handed.
The worker pool the spawns are dispatched to is no longer odls's, so its
sizing and rotation are covered in test/unit/runtime/test_runtime.c
(test_worker_pool) rather than here.
Add structural regression guards here; anything that needs a running launch belongs in the swarm harness or the integration harness.
Where to go next
The launch primitive itself lives in the component:
pdefault/AGENTS.md— the default (and only) local-launch component: the realfork()/execve(), the parent/child pipe protocol, and how it plugsfork_local_procinto the base.
