vignettes/articles/cli-esgf-store.Rmd
cli-esgf-store.RmdThe epwshiftr command line interface is a thin wrapper around the package’s R APIs. It is meant for day-to-day ESGF store operations and future-weather workflows: checking the local environment, creating and tracking queries, previewing updates, running downloads, validating local files, extracting site climate data, morphing baseline EPW files, and writing future EPW outputs.
The CLI mirrors the same layers documented in ESGF query results, ESG stores, Downloader, Create Future EPW Files, and ESGF troubleshooting. Workflow
state is store backed. Commands exchange persistent
query_id, plan_id, and morph_id
values from the EsgStore manifest rather than writing
separate stage files.
It is not a separate cross-platform binary. The launcher installed by
install_cli() calls the current Rscript and
runs epwshiftr::epwshiftr_cli(exit = TRUE). This keeps
command-line behaviour close to the R package while still giving you a
convenient epwshiftr command.
Install the launcher once from R:
On macOS and Linux, the default location is
~/.local/bin/epwshiftr. On Windows, it is
%LOCALAPPDATA%/epwshiftr/bin/epwshiftr.cmd.
install_cli() does not edit shell profiles or
PATH; if the target directory is not on PATH,
add it yourself or call the launcher by its full path.
You can also use the CLI entry point directly without installing a launcher:
Use doctor before longer runs. The default checks are
local and read-only. Add --network only when you want to
probe an ESGF index node.
Most commands operate on an EsgStore. Use
--store to choose the store for one command:
When --store is omitted, the CLI uses
store_dir(), the same default store location used by the R
API.
Use shift run --config when you want one command to
execute the main workflow: request, collect, optional download, extract,
morph, and write EPW files. The configuration file is JSON and is
validated before work starts. The transform selects the scientific
calculation; run-specific model and observed references are separate
top-level inputs. Original morphing requires a matching historical model
reference.
{
"version": 2,
"epw": "baseline/SIN.epw",
"climate": {
"provider": "cmip6",
"model": "BCC-CSM2-MR",
"scenarios": ["ssp126", "ssp585"],
"member": null,
"grid": null,
"frequency": "mon",
"table": null
},
"periods": {
"2060s": "2055:2065"
},
"transform": {
"scale": "monthly",
"method": "original_morphing"
},
"reference": {
"mode": "historical",
"periods": {
"reference": "1995:2014"
}
},
"dir": "outputs/future-epw",
"control": {
"strict": true,
"allow_partial": false,
"download": "auto",
"resume": true,
"overwrite": false
}
}climate.table = null enables transform-aware table
selection. A scalar forces every variable into one table; a JSON object
such as {"snd": "LImon"} overrides only the named variable.
The resolved transform and its scientific options are persisted with the
run. Version 1 workflow records are rejected rather than silently
reinterpreted under the new API.
Preview the parsed request and site without writing store state:
Run the workflow. control.download owns the access
policy; "auto" prefers OPeNDAP and falls back to HTTP
downloads.
For a detached worker, add --background. Progress
presentation is a CLI concern rather than part of the JSON scientific
configuration:
The result includes one durable run ID. All workflow inspection and recovery commands use that ID, so they do not need users to coordinate query, extraction, and morph identifiers manually:
epwshiftr --store ~/cmip6-store shift status --run <run_id>
epwshiftr --store ~/cmip6-store shift show --run <run_id>
epwshiftr --store ~/cmip6-store shift watch --run <run_id> --follow
epwshiftr --store ~/cmip6-store shift diagnostics --run <run_id>
epwshiftr --store ~/cmip6-store shift outputs --run <run_id>
epwshiftr --store ~/cmip6-store shift data --run <run_id> --limit 5
epwshiftr --store ~/cmip6-store shift logs --run <run_id> --tail 100
epwshiftr --store ~/cmip6-store shift cancel --run <run_id>
epwshiftr --store ~/cmip6-store shift resume --run <run_id> --backgroundCtrl-C during shift watch --follow stops
watching without cancelling the worker. A normal
shift cancel stops at the next safe workflow boundary. Use
shift cancel --force only when the worker does not respond
to that request.
You can generate and validate workflow JSON from the CLI:
The method catalog exposes every selectable scale and reconstruction, including experimental status, evidence type, input roles, and output type. These commands work offline without opening a store:
epwshiftr morph transforms
epwshiftr morph transforms --scale daily --status experimental
epwshiftr morph describe --scale hourly --method kernel_qdm
epwshiftr --json morph describe --scale daily --method qdm \
--option 'tas.bounds=[-40,60]'morph describe shows effective settings alongside
constructor defaults. It checks candidate settings with the same
validator used for execution. Quote JSON vectors and objects so the
shell preserves them; repeat --option for different
settings. Field roles distinguish transformed, derived, physically
closed, and inherited EPW fields. Experimental status is a method
support boundary, and does not by itself establish field validation.
Generate a method/model batch and its required ERA5 calibration input:
epwshiftr shift config example --methods original_morphing,qdm \
--model 3 --output batch.json
epwshiftr shift config validate --config batch.json
epwshiftr --store ~/cmip6-store shift config validate \
--config batch.json --network
epwshiftr doctor --config batch.json
epwshiftr --store ~/cmip6-store shift run --config batch.json --dry-run
epwshiftr --store ~/cmip6-store shift run --config batch.json --backgroundEdit the generated baseline EPW path, periods, output directory, and
calibration years for your study. climate.model accepts a
positive count, explicit model names, or null for all
compatible models (--model all in the example command).
climate.frequency: null lets each method select its source
resolution. The calibration object accepts
dataset, years, product,
variables, frequency, access, and
non-secret request options. ERA5 credentials stay in the
existing CDS configuration or environment, outside workflow JSON.
observed_reference remains supported as an alternative to
calibration; do not supply both. ERA6 specifications
currently report an explicit unavailable provider error.
Local validation checks schema, options, local EPW readability, and
reference roles. It does not establish remote coverage.
--network checks common model coverage and reanalysis
readiness; shift run --dry-run also resolves the model
matrix. These operations may contact providers but do not generate EPWs.
Their human output summarizes the baseline, requested methods and
models, future periods, historical reference, calibration input, and
destination. Network checks and discovery show the current method, node,
and variables. Use --no-progress to retain the final report
without live updates; global --quiet, --json,
and --jsonl also suppress human progress output.
Method keys with several possible configurations need explicit
objects. For example, replace methods with this
transform array to select both daily reconstructions:
"transform": [
{"scale": "daily", "method": "epwshiftr", "reconstruction": "power"},
{"scale": "daily", "method": "epwshiftr", "reconstruction": "btws"}
]A batch returns a durable batch_id. Use exactly one of
--batch and --run:
epwshiftr --store ~/cmip6-store shift show --batch <batch_id>
epwshiftr --store ~/cmip6-store shift show --batch <batch_id> --verbose
epwshiftr --store ~/cmip6-store shift watch --batch <batch_id> --follow
epwshiftr --store ~/cmip6-store --jsonl shift watch --batch <batch_id> --follow
epwshiftr --store ~/cmip6-store --json shift outputs --batch <batch_id>
epwshiftr --store ~/cmip6-store shift diagnostics --batch <batch_id>
epwshiftr --store ~/cmip6-store shift resume --batch <batch_id> --background
epwshiftr --store ~/cmip6-store shift cancel --batch <batch_id>New dry-run receipts persist child plans, so
resume --batch can start them in a later session without
repeating discovery. In R, reopen the same receipt with
shift_batch_get(batch_id, store) and use the ordinary
shift_*() inspectors. Older receipts containing child runs
remain readable; older dry-run receipts without saved plans require the
original configuration to be run again.
The batch dashboard uses the same bordered panel as a single
workflow, with separate overview, Workflows, and
Results sections. Each child shows its method/model
identity above its reconstruction and stage; narrow terminals retain
these rows without borders. A failed child does not stop watching
another active child. Completion summaries distinguish cases from
physical EPW files: a multi-year output can write several files for one
case. Output rows retain method/model identity, weather year, output
type, and field provenance. Warnings remain available in diagnostics and
appear in completion summaries.
Dynamic batch views fit the terminal height, put failed and active
children first, and show the current operation, measured progress, and
last update. Cancelled children have their own count. Errors precede
warnings and retain recovery actions. Use
shift show --verbose for all children and diagnostics; a
completed or count-limited watch leaves its final receipt in scrollback.
For a single run, shift show --run <run_id> --verbose
expands every case, output, event, and diagnostic into wrapped records
with complete paths and recovery actions. Add --debug to
include raw JSON payloads. These flags only affect human-readable
output; JSON and JSONL keep their existing fields. Batch watch follows
each child’s event history independently, including late events from
parallel workers and events received with the final snapshot.
Find previous work and summarize its outputs without contacting providers:
epwshiftr --store ~/cmip6-store shift list
epwshiftr --store ~/cmip6-store shift list --type batch --status failed,partial
epwshiftr --store ~/cmip6-store shift summary --batch <batch_id>
epwshiftr --store ~/cmip6-store --json shift summary --run <run_id> --weathershift list returns the newest 20 records by default
(--limit N changes this) with complete IDs and store paths.
Batch children show their parent batch and own store; use that child
store when passing its run ID to another command. An unreadable record
remains visible as unavailable with an error.
shift summary groups cases and outputs by method, model,
scenario, period, member, and grid. It distinguishes recorded EPWs from
available local files and retains output type, weather years,
diagnostics, and field roles. --weather reads local EPWs
and adds mean temperature, relative humidity, wind speed, and global
horizontal radiation, with a valid-hour count for each mean and explicit
file-read errors. Missing EPW sentinel values are excluded. Check the
output type, years, and hourly coverage before comparing methods; the
means do not rank methods or make unlike output periods equivalent.
Repeated completed batch calls reuse their persisted receipts and
verified artifacts. The human receipt reports reused versus
started/resumed children and the elapsed time of the current call. Set
control.refresh: true to explicitly refresh catalog
selection. CLI execution returns status 1 for failed,
blocked, partial, or cancelled results, and status 2 for
invalid command or configuration syntax; inspection commands can still
read those results.
Use extract commands when a query already exists in the
store and you want to plan, run, and inspect the climate extraction
stage directly.
epwshiftr --store ~/cmip6-store extract plan \
--query <query_id> \
--site-id SIN \
--lon 103.98 \
--lat 1.37 \
--time 2055-01-01T00:00:00Z,2064-12-31T23:59:59Z \
--variable tas,hurs,psl,rsds,rlds,sfcWind,clt \
--nearest 1
epwshiftr --store ~/cmip6-store extract run --plan <plan_id> --fallback auto
epwshiftr --store ~/cmip6-store extract retry --plan <plan_id>
epwshiftr --store ~/cmip6-store extract retry --plan <plan_id> --run
epwshiftr --store ~/cmip6-store extract coverage --plan <plan_id>
epwshiftr --store ~/cmip6-store extract artifacts --plan <plan_id>Use morph commands to inspect available transforms, run
a transform from extracted plan_id values, and write EPW
outputs later from the stored morph_id.
epwshiftr --store ~/cmip6-store morph transforms
epwshiftr --store ~/cmip6-store morph variables \
--scale monthly \
--method original_morphing
epwshiftr --store ~/cmip6-store morph run \
--plan <plan_id> \
--epw baseline/SIN.epw \
--period 2060s=2055:2064 \
--scale monthly \
--method original_morphing \
--reference historical \
--reference-period reference=1995 \
--strict true
epwshiftr --store ~/cmip6-store morph epw --morph <morph_id> --dir outputs/future-epw
epwshiftr --store ~/cmip6-store morph retry --morph <morph_id>
epwshiftr --store ~/cmip6-store morph retry --morph <morph_id> --run
epwshiftr --store ~/cmip6-store morph status --morph <morph_id>
epwshiftr --store ~/cmip6-store morph outputs --morph <morph_id>Repeat --option KEY=VALUE for each scientific setting
supported by the selected transform. Use --reconstruction
only when morph transforms lists more than one
reconstruction for that scale and method.
Variable-specific settings use VARIABLE.SETTING=VALUE;
JSON arrays preserve vector settings:
query search builds an EsgQuery from
key=value constraints. The command contacts ESGF unless you
pass --dry-run.
epwshiftr query search --dry-run \
--index-node https://esgf-data.dkrz.de \
--type Dataset \
project=CMIP6 \
activity_id=ScenarioMIP \
experiment_id=ssp585 \
variable_id=tas,psl \
table_id=Amon \
datetime_start=2050-01-01 \
datetime_stop=2050-12-31 \
latest=trueComma-separated values become multi-value query parameters. Boolean
strings true and false become logical query
values.
Use --fields to control fields requested from ESGF, and
--columns to control the human-readable table printed by
the CLI. If a display column is an ESGF result field, include it in
--fields as well.
epwshiftr query search \
--fields source_id,experiment_id,variable_id,datetime_start,datetime_stop,size \
--columns variable_id,datetime_start,datetime_stop,size \
project=CMIP6 \
variable_id=tas \
datetime_start=2050-01-01 \
datetime_stop=2050-12-31Use datetime_start and datetime_stop for
data coverage constraints. These are mapped to
EsgQuery$datetime_range(), so they are not plain ESGF
facets. For a target year, set both boundaries:
epwshiftr query search --dry-run \
project=CMIP6 \
variable_id=tas \
datetime_start=2050-01-01 \
datetime_stop=2050-12-31datetime_end is accepted as an alias for
datetime_stop.
Use key!=value for a negated facet constraint:
A single key cannot mix positive and negated constraints in the same
command. Use either source_id=A,B or
source_id!=A,B, not both.
Use query add when a search should become part of a
persistent store. This saves the query definition in the store; it does
not download files.
epwshiftr --store ~/cmip6-store query add \
--label singapore-ssp585-amon \
--tag cmip6 \
--tag future-weather \
--track \
--index-node https://esgf-data.dkrz.de \
project=CMIP6 \
activity_id=ScenarioMIP \
experiment_id=ssp585 \
variable_id=tas,psl \
table_id=Amon \
datetime_start=2050-01-01 \
datetime_stop=2050-12-31 \
latest=trueUse --dry-run to inspect the query URL, label, tags, and
tracked flag before writing to the store.
epwshiftr --store ~/cmip6-store query add --dry-run \
--label singapore-ssp585-amon \
project=CMIP6 variable_id=tas table_id=AmonYou can also import a query JSON file saved from R:
Tracked queries support a two-step update path. First preview the diff without changing the store:
epwshiftr --store ~/cmip6-store query preview <query_id>
epwshiftr --store ~/cmip6-store query preview <query_id> --detailThen persist the update when the diff is acceptable:
Inspect query metadata, files, update batches, and non-current changes:
download preflight refreshes query results, computes the
update diff, expands replica candidates, and estimates the download set
without enqueueing tasks or downloading data.
epwshiftr --store ~/cmip6-store download preflight <query_id> \
--replica auto \
--service HTTPServer \
--strategy fastestFor a conservative dry check, disable network probing:
download run performs the full store-aware path:
refresh, plan, enqueue, run, and sync completed files back into the
store catalog.
epwshiftr --store ~/cmip6-store download run <query_id> \
--session-label singapore-ssp585-amon \
--replica auto \
--service HTTPServerFor long downloads, run in the background.
--mode process starts a detached R process for this job.
--mode daemon submits the job to a running downloader
daemon.
epwshiftr --store ~/cmip6-store download run <query_id> \
--session-label singapore-ssp585-amon \
--background \
--mode processThe downloader keeps persistent sessions, tasks, candidate URLs, data-node statistics, jobs, logs, and events. You can stop R and later inspect or resume the same session.
epwshiftr --store ~/cmip6-store download sessions
epwshiftr --store ~/cmip6-store download tasks --session <session_id>
epwshiftr --store ~/cmip6-store download watch --query <query_id>
epwshiftr --store ~/cmip6-store download logs --session <session_id> --tail 50
epwshiftr --store ~/cmip6-store download jobs
epwshiftr --store ~/cmip6-store download logs --job <job_id> --tail 50
epwshiftr --store ~/cmip6-store download stop --job <job_id>Use daemon mode when you want one persistent process to run queued jobs over time:
epwshiftr --store ~/cmip6-store download daemon start
epwshiftr --store ~/cmip6-store download run <query_id> --background --mode daemon
epwshiftr --store ~/cmip6-store download daemon status
epwshiftr --store ~/cmip6-store download daemon stopRetry failed or cancelled tasks:
epwshiftr --store ~/cmip6-store download retry --query <query_id>
epwshiftr --store ~/cmip6-store download retry --query <query_id> --runResume interrupted tasks or verify completed files:
Show the current persistent downloader configuration:
Set operational limits when running large batches:
epwshiftr --store ~/cmip6-store download config set \
--workers 4 \
--host-concurrency 2 \
--timeout 120 \
--chunk-size 1048576 \
--bandwidth-limit noneInspect data-node performance records and reset stale node statistics if needed:
The storage commands expose the store’s file-system maintenance tools.
Choose a download layout policy:
epwshiftr --store ~/cmip6-store storage layout show
epwshiftr --store ~/cmip6-store storage layout set --layout drsSummarise disk usage and validate cataloged files:
epwshiftr --store ~/cmip6-store storage report
epwshiftr --store ~/cmip6-store storage report --detail
epwshiftr --store ~/cmip6-store storage validate --query <query_id>Checksum validation can be expensive on large NetCDF collections, so it is opt-in:
Repairs and cleanup commands default to preview mode. Add
--execute only after reviewing the proposed actions.
Human-readable output is the default. Add --json when
another script should consume a complete result. Add
--jsonl with download watch --follow when a
script should consume streaming progress events one object per line.
epwshiftr --store ~/cmip6-store --json query status <query_id>
epwshiftr --store ~/cmip6-store --json download watch --query <query_id>
epwshiftr --store ~/cmip6-store --json storage validate --query <query_id>
epwshiftr --store ~/cmip6-store --jsonl download watch --query <query_id> --followUse --quiet when you only need the process status. CLI
status codes are:
0: success;1: runtime error;2: usage error.Use built-in help for the complete command list:
epwshiftr help
epwshiftr help query add
epwshiftr help download run
epwshiftr help shift run
epwshiftr help extract plan
epwshiftr help morph run
epwshiftr help storage validateThe command groups intentionally mirror the package boundaries:
doctor: local and optional network diagnostics;query: create, track, update, inspect, tag, and remove
store queries;download: preflight, run, resume, retry, verify, and
inspect downloads;shift: config-driven request, collect, extract, morph,
and EPW workflow runs;extract: store-ID extraction planning, execution,
coverage, and artifacts;morph: transformation variables, method catalog, runs,
EPW writing, status, and outputs;storage: layout, report, validate, repair, and cleanup
local files;esgf: compact ESGF query and download health
reports.