Skip to main content

sacct_publish

Overview​

Transforms and publishes sacct output in parsable2 format to configured sinks. Takes raw sacct command output, parses pipe-delimited (parsable2) format, converts timestamps to timezone-aware ISO8601 format, enriches with cluster metadata, and publishes to sink in batches with retry logic.

Data Type: DataType.LOG Schema: Sacct

Execution Scope​

Invoked as a pipeline component (typically by sacct_backfill or in shell pipelines).

Output Schema​

Sacct​

Published with DataType.LOG:

{
# SLURM Fields (from sacct -o all)
"JobID": str, # Job identifier
"JobIDRaw": str, # Raw job ID
"JobName": str, # User-specified job name
"User": str, # Job owner
"Account": str, # SLURM account
"Partition": str, # Job partition
"State": str, # Job state (COMPLETED, FAILED, etc.)
"ExitCode": str, # Exit code
"AllocCPUs": int, # Allocated CPUs
"ReqMem": str, # Requested memory
"AllocNodes": int, # Allocated nodes
# ... plus 40+ additional sacct fields

# Timestamps (converted to ISO8601 with timezone)
"Start": str, # Job start time
"End": str, # Job end time
"Submit": str, # Submission time
"Eligible": str, # Eligible time

# Enriched Fields
"cluster": str, # Cluster name
"derived_cluster": str, # Sub-cluster (same as cluster if not `--heterogeneous-cluster-v1`)
"sc_host": str | None, # Kubernetes host for a single-node job
"sc_node_hosts": str | None, # JSON Slurm-node-to-Kubernetes-node map
"time": int, # Collection timestamp (Unix epoch)
"end_ds": str, # Date string in Pacific Time (YYYY-MM-DD)
}

Note: The exact fields depend on the SLURM version. See sacct documentation for details.

Command-Line Options​

OptionTypeDefaultDescription
SACCT_OUTPUTFile/stdinFile or stdin containing sacct output
--clusterStringAuto-detectedCluster name for metadata enrichment
--sinkStringRequiredSink destination, see Exporters
--sink-optsMultiple-Sink-specific options
--log-levelChoiceINFODEBUG, INFO, WARNING, ERROR, CRITICAL
--log-folderString/var/log/fb-monitoringLog directory
--stdoutFlagFalseDisplay metrics to stdout in addition to logs
--heterogeneous-cluster-v1FlagFalseEnable per-partition metrics for heterogeneous clusters
--intervalInteger300Seconds between collection cycles (5 minutes)
--onceFlagFalseRun once and exit (no continuous monitoring)
--retriesIntegerShared defaultRetry attempts on sink failures
--dry-runFlagFalsePrint to stdout instead of publishing to sink
--chunk-sizeIntegerShared defaultThe maximum size in bytes of each chunk when writing data to sink.
--delimiterString|#Field delimiter, must be unique enough such that no field has an instance of the delimiter. User input fields like job name can break parsing.
--sacct-timezoneStringSystem timezoneTimezone of sacct timestamps
--sacct-output-io-errorsChoicestrictUTF-8 decoding: strict, ignore, replace. Occasionally, user-defined sacct fields (e.g. comment, job name) are not valid utf-8, so we allow the user to choose what should be done in this case
--ignore-line-errorsFlagFalseSkip lines with field count mismatches
--resolve-kubernetes-hostsFlagFalseResolve single-node Slurm pod names to Kubernetes host names
--kubernetes-namespaceStringRequired when enabledKubernetes namespace containing Slurm compute pods
--kubernetes-label-selectorStringAll podsLabel selector identifying Slurm compute pods

Usage Examples​

From stdin (piped from sacct)​

sacct -P -o all -S 2024-01-01 -E 2024-01-02 | \
gcm sacct_publish --sink stdout

From file​

gcm sacct_publish sacct_output.txt \
--sink graph_api \
--sink-opts scribe_category=slurm_jobs

With timezone handling​

# sacct output is in UTC
gcm sacct_publish sacct_output.txt \
--sacct-timezone UTC \
--sink file \
--sink-opts filepath=output.json

Ignore invalid lines​

# Skip lines with malformed data
gcm sacct_publish sacct.txt \
--ignore-line-errors \
--sacct-output-io-errors replace \
--sink stdout

With sacct_backfill​

gcm sacct_backfill new \
--publish-cmd "gcm sacct_publish --sink graph_api --sink-opts scribe_category=jobs"

Kubernetes host enrichment​

sacct_publish can resolve the Slurm compute pods in sacct.NodeList to their Kubernetes nodes. It publishes every resolved pair as a JSON object in the string column sc_node_hosts. When NodeList contains exactly one node, it also publishes that node as sc_host:

gcm sacct_publish sacct.txt \
--resolve-kubernetes-hosts \
--kubernetes-namespace tenant-slurm \
--kubernetes-label-selector app.kubernetes.io/component=compute \
--sink graph_api \
--sink-opts scribe_category=perfpipe_fair_sacct

For example, g3-130-[015,021] may produce:

{"g3-130-015": "sf8gg2h4", "g3-130-021": "sf8gg2h9"}

This requires the kubernetes optional dependency and read-only permission to list pods in the configured namespace. Multi-node records leave sc_host unset so that the scalar value is never ambiguous. Missing pods are omitted from the JSON map. Kubernetes client or API failures time out after five seconds and do not prevent the original sacct records from being published. Because resolution uses a current pod snapshot, deleted historical pods cannot be enriched.