sliderule-runner is a command line tool for running a SlideRule script over a large list of inputs (for example, a list of granules) as a set of batch jobs. It can be used to submit jobs, monitor their progress, summarize their results, collect the inputs that failed so they can be resubmitted, retrieve job logs, and cancel jobs. The tool keeps a local database of everything you have submitted so that you can come back to your jobs later.
Quick Start¶
# Run script.lua once for every line in granules.txt
sliderule-runner submit my_job script.lua granules.txt
# Check on progress
sliderule-runner status
# Summarize the results of completed jobs
sliderule-runner report
# Save the entries that did not succeed so they can be resubmitted
sliderule-runner scrape --output retry.txt
sliderule-runner submit my_job_retry script.lua retry.txtwhere granules.txt contains one entry per line:
ATL03_20230601000000_10011901_006_01.h5
ATL03_20230601001000_10021901_006_01.h5
ATL03_20230602000000_10031901_006_01.h5granules.txt
Usage¶
sliderule-runner <command> [options]Run sliderule-runner --help for a list of commands, or sliderule-runner <command> --help for the options of a specific command.
Common Options¶
The following options are accepted by every command. They are attached to each subcommand, so they must appear after the command name:
Option | Default | Description |
|---|---|---|
|
| Domain of the SlideRule service to connect to. |
|
| Location of the local database that tracks your submissions (see The Job Database). |
|
| Priority queue that jobs are submitted to and looked up in. Valid values are |
| False | Turn on verbose log messages. Also shows more detail in |
| False | Do not save changes to the local database when the command finishes. |
Command Reference¶
submit¶
sliderule-runner submit <name> <script.lua> <arguments.txt> [options] [common options]Submits a job that runs a script once for each entry in an arguments file. If the arguments file has more entries than --batch_size, the entries are split into multiple jobs, each of which is submitted and recorded separately.
Argument / Option | Default | Description |
|---|---|---|
| (required) | Base name for the submission. See Job names. |
| (required) | Path to the script to run for each entry. |
| (required) | Path to a file containing one entry per line. With |
| False | Treat the arguments positional as a single string entry rather than the name of a file. |
|
| Maximum number of entries per job. Larger arguments files are split into multiple jobs. |
|
| Number of virtual CPUs allocated to the job. |
|
| Memory in KBytes allocated to the job. (For example, 16000 is 16MB) |
|
| Container image used to run the job. |
Job names¶
Each job is named <name>_<i>, where <i> is the position in the arguments file of the first entry in that batch. For example, submitting my_job with 25,000 entries and the default batch size creates my_job_0, my_job_10000 and my_job_20000. If a name is already in your local database, three random letters are added (e.g. my_job_xqf_0) so that existing records are not overwritten.
Examples:
# run script.lua on every entry in granules.txt
sliderule-runner submit my_job script.lua granules.txt
# split the entries into jobs of 500 and give each job more resources
sliderule-runner submit my_job script.lua granules.txt --batch_size 500 --vcpus 8 --memory 32000
# run the script on a single entry supplied on the command line
sliderule-runner submit one_off script.lua "ATL03_20230601000000_10011901_006_01.h5" --arg_as_str
# submit to a different queue using a custom image
sliderule-runner submit my_job script.lua granules.txt --queue <priority> --image sliderule:my-branchstatus¶
sliderule-runner status [common options]Checks on every submission in the local database that is not yet complete, updates the database, and prints a table of the current state of every submission. When all of the work in a submission has finished, the tool reads its results from S3 and marks the submission as complete, after which it is available to report and scrape.
Because the tool does not run in the background, status needs to be run again to pick up progress.
By default, the table has one row per submission and one column for each job state:
Column | Description |
|---|---|
| The full job name. |
| Number of jobs in each in-progress state. |
| Number of jobs that have finished in each state. |
With --verbose, the table instead has one row per child job, showing the submission NAME, the child job INDEX and its STATUS. This index is the value to give to logs --index.
Examples:
# check on all submitted jobs
sliderule-runner status
# show the status of each individual job in each submission
sliderule-runner status --verbose
# look at the status without saving any updates to the local database
sliderule-runner status --dryrunreport¶
sliderule-runner report [--name <str>] [common options]Generates a summary of the results of completed submissions. Submissions that are not yet complete are not included.
Option | Default | Description |
|---|---|---|
| none | Only report on the submission with this full job name. |
The report has one row per completed submission with the following columns:
Column | Description |
|---|---|
| The full job name. |
| Number of entries that have a recorded duration. |
| Reserved; currently always 0. |
| Number of entries with each result status (see Job Results). |
| Average duration per processed entry, in seconds. |
Examples:
# report on all completed submissions
sliderule-runner report
# report on a single submission
sliderule-runner report --name my_job_0scrape¶
sliderule-runner scrape [options] [common options]Generates a list of the original entries whose results have one of the given statuses. By default it selects the entries that did not succeed, which makes it easy to build the arguments file for a resubmission. Submissions that are not yet complete are skipped.
Option | Default | Description |
|---|---|---|
|
| One or more result statuses to select (see Job Results). |
| none | Only scrape the submission with this full job name. |
| none | Write the selected entries to this file, one per line. |
Each selected entry is printed with its position in the submission. With --verbose, the complete result record for each entry is printed (and written to --output) instead of just the entry.
Examples:
# list everything that did not succeed across all completed submissions
sliderule-runner scrape
# save the entries that did not succeed to a file, ready to resubmit
sliderule-runner scrape --output retry.txt
# save only the entries that failed in a single submission
sliderule-runner scrape --name my_job_0 --status failure --output retry.txtlogs¶
sliderule-runner logs --name <str> [--index <int>] [common options]Prints the log messages for a job. Each message is printed on its own line, preceded by ->.
Option | Default | Description |
|---|---|---|
| (required) | Full job name of the submission. |
| none | Index of a single child job within the submission, as shown by |
Examples:
# find the index of the job you are interested in
sliderule-runner status --verbose
# get the logs for child job 17 of the submission
sliderule-runner logs --name my_job_0 --index 17cancel¶
sliderule-runner cancel --name <str> [common options]Cancels a submitted job. For a submission made up of many entries, this cancels the submission as a whole.
Option | Default | Description |
|---|---|---|
| (required) | Full job name of the submission to cancel. |
Examples:
sliderule-runner cancel --name my_job_0archive¶
sliderule-runner archive <full path to archive file> [common options]Saves the contents of the local database to an archive file and then clears the database. Use it to start fresh once you are finished with a set of jobs.
Argument | Description |
|---|---|
| (required) Where to save a copy of the database before it is cleared. |
Examples:
sliderule-runner archive ~/sliderule_archives/2026-09-runs.jsonAdditional Topics¶
The Job Database¶
The tool keeps a JSON database of your submissions at ~/.cache/sliderule/runner_database.json (change it with --database). For each submission it records what was returned when the job was submitted, the number of entries, the latest status, the child jobs, and, once the submission is complete, the results for every entry.
The database is written when a command finishes successfully. If a command fails part way through, none of its changes are saved.
Use
--dryrunto run a command without saving changes to the database.Use
archiveto save a copy and start with an empty database.Use
--databaseto keep separate sets of work in separate databases.
Job Results¶
When a submission finishes, the tool reads the output of each entry from S3. For this to work, the result your script writes for each entry must be a JSON object with the following fields:
{
"status": true,
"start": 1719849600.0,
"stop": 1719849660.0,
"outputs": ["<output1>", "<output2>"]
}Field | Description |
|---|---|
| Boolean: whether the run for this entry succeeded. |
| Start time of the run, in seconds. |
| Stop time of the run, in seconds. |
| List of the outputs produced for this entry. |
Each entry is then given one of the following result statuses, which are the values used by report and by scrape --status:
Status | Meaning |
|---|---|
| The result was read and its |
| The result was read and its |
| The result was read, but it does not have the fields listed above, so it could not be interpreted. |
| The result could not be read. |
Typical Workflow¶
# 1. submit
sliderule-runner submit my_job script.lua granules.txt
# 2. monitor until everything is complete (repeat as needed)
sliderule-runner status
# 3. summarize
sliderule-runner report
# 4. investigate anything that went wrong
sliderule-runner status --verbose
sliderule-runner logs --name my_job_0 --index 17
# 5. collect what did not succeed and resubmit it
sliderule-runner scrape --output retry.txt
sliderule-runner submit my_job_retry script.lua retry.txt
# 6. when finished, archive the database
sliderule-runner archive ~/sliderule_archives/my_job.jsonOutput Format¶
status, report and scrape print their results to standard output. status and report print tables of comma-separated, right-aligned columns with a header row. Progress bars are shown while results are being read if the optional tqdm package is installed; if it is not, a message is printed when the tool starts and progress is not reported.
Error Handling¶
By default, errors are caught and reported as a single line followed by a traceback:
Unhandled error: <message>Pass --verbose to raise the exception directly instead. When a command fails, the local database is not updated.