Data Ingestion
Data Archive
When working with local data, you must first download it from the Subaru data archive for your open-use program, specifically from the STARS2 database. The archive provides access to science images, calibration data, and PfsConfig files.
Note
Details on data retrieval will be provided by the STARS team on the day following your observation.
Typically, you will download a TAR file from the STARS2 website. After that, follow the steps below to extract and access the data on UNIX or macOS:
# Step 1: Create a directory for downloads
mkdir $WORKDIR/$(whoami)/download_dir
# Step 2: Copy the downloaded TAR file to the directory
cp S2Query.tar $WORKDIR/$(whoami)/download_dir
# Step 3: Navigate to the directory
cd $WORKDIR/$(whoami)/download_dir
# Step 4: Extract the TAR file
tar -xvf S2Query.tar
# Step 5: Run the unpacking script
./zadmin/unpack.py
Platform-specific instructions:
- UNIX: Use
wgetfor downloading. - macOS: Use
curlfor downloading. - Windows: An FTP manager is required for data transfer.
Ingestion to butler
You can ingest data into the butler repository. You will need the custom pfsConfig files created for your observing program. The custom pfsConfig files can be found on the Science Platform under your proposal ID as follows:
/shared/pfs/programs/$PROPOSAL_ID/2d/customPfsConfig/
Copy these files to your local disk and perform the steps below to ingest them.
There are two types of data that need to be ingested: raw images and PfsConfig files.
These are ingested using two separate commands:
# Assume that we define the destination datastore (`DATASTORE`) and the directory containing the input data (`DATADIR`):
DATASTORE="$WORKDIR/$(whoami)/data/datastore"
DATADIR="$WORKDIR/$(whoami)/data"
# Ingest the raw images taken in March 2026
butler ingest-raws $DATASTORE $DATADIR/raw/2026-03-*/*/PFS*.fits --ingest-task lsst.obs.pfs.gen3.PfsRawIngestTask --transfer link --fail-fast
# Ingest the PfsConfig files for the March 2026 run
ingestPfsConfig.py $DATASTORE PFS PFS/raw/pfsConfig $DATADIR/raw/2026-03-*/pfsConfig/pfsConfig-*.fits --transfer link
If you are using PFSF files provided by the observatory, they are equivalent to pfsConfig files, and a similar command can be used:
ingestPfsConfig.py $DATASTORE PFS PFS/raw/pfsConfig $DATADIR/raw/2026-03-*/pfsConfig/PFSF*.fits --transfer link
For details on the ingested data, refer to the Appendix or the datamodel.
For example, the filenames have the following meanings:
- PFSA12345611.fits:
A raw
scienceexposure withvisit=123456, taken atsite=summitwithspectrograph=1using the blue arm (armNum=1). - PFSB12345623.fits:
A raw
up-the-rampexposure withvisit=123456, taken atsite=summitwithspectrograph=2using the IR arm (armNum=3). - pfsConfig-0xad349fe21234abcd-123456.fits:
A realization of a
PfsDesignwithpfsDesignId=ad349fe21234abcdforvisit=123456. - PFSF12345600.fits:
A realization of a
PfsDesignused by the observatory forvisit=123456. In the final two digits,00indicates the original fullPfsConfig(PFSF) file;01-99indicate customizedPFSFfiles containing fibers associated with a specific proposal ID together with calibration fibers (for example, sky and flux fibers). The observatory will provide the customizedPFSFfiles with01-99in the final two digits. If thePFSFfile contains only one proposal ID or one calibration frame, the00file will be distributed. The ingestion procedure is the same as for aPfsConfigfile.
The parameters in the commands include:
--transfer: The method used to add data to the repository. The options includelink,copy, andmove, which specify whether the data is symlinked, duplicated, or physically relocated, respectively.--fail-fast: Stops the ingestion process immediately if an error occurs. This is useful for debugging. If you do not need this behavior, omit this option.
The ingestion process places the files (referred to as "datasets" in butler) in the repository and records them in the registry database. Each file is placed in a collection, which can be thought of as a directory-like grouping in the butler (and, when using a traditional filesystem datastore, it is implemented as a directory).
The raw data are placed in the collection PFS/raw/sps, while the PfsConfig files are placed in the collection PFS/raw/pfsConfig.
Troubleshooting
When ingesting data into DATASTORE, all files for a single visit must be ingested with a single command. For example, the following commands ingest two files from visit=123456 separately and will fail:
# The second command will fail
butler ingest-raws $DATASTORE $DATADIR/raw/2026-03-18/sps/PFSA12345601.fits --ingest-task lsst.obs.pfs.gen3.PfsRawIngestTask --transfer link --fail-fast
butler ingest-raws $DATASTORE $DATADIR/raw/2026-03-18/sps/PFSA12345602.fits --ingest-task lsst.obs.pfs.gen3.PfsRawIngestTask --transfer link --fail-fast
The correct approach is to use wildcards so that all files from visit=123456 are included at once:
butler ingest-raws $DATASTORE $DATADIR/raw/2026-03-18/*/PFS*123456*.fits --ingest-task lsst.obs.pfs.gen3.PfsRawIngestTask --transfer link --fail-fast
To re-ingest one or more visits, first prune the previously ingested data, then repeat the procedure above:
butler prune-datasets $DATASTORE PFS/raw/sps --datasets=raw --unstore --where="instrument='PFS' AND visit=123456"
# Additional dimensions can be specified by using an SQL expression
butler prune-datasets $DATASTORE PFS/raw/sps --datasets=raw --unstore --where="instrument='PFS' AND visit IN (123456,123457) AND arm='n' AND spectrograph='3'"
Dataset
Each dataset is specified by a dataId, which is a dictionary of key-value pairs representing the dimensions.
For example,
- A
rawimage may have adataIdsuch as {'instrument': 'PFS', 'visit':123, 'arm': 'r', 'spectrograph':3}. - A
PfsConfigfile is valid for an entire exposure, so it may have adataIdsuch as {'instrument': 'PFS', 'visit':123}.
IMPORTANT: In general, users should treat the files in the datastore as a butler implementation detail, and use the butler commands and Python API to access the data products.
There are some kinds of datastores that do not use a traditional filesystem (e.g., the S3 datastore), and so the files may not be directly accessible.
Warning
The registry database tracks all files in the datastore. Do not delete files from the datastore without using the appropriate butler commands.
You can see what raw datasets are in the datastore with the following command:
butler query-datasets $DATASTORE --collections PFS/raw/sps
The result looks something like this:
type run id instrument arm dither pfs_design_id spectrograph detector visit
---- ------------- ------------------------------------ ---------- --- ------ ------------- ------------ -------- --------
raw PFS/raw/all 27217522-a357-5071-a32b-af97b5b8bee6 PFS b 0.0 1 1 0 0
raw PFS/raw/all 0ce0cbea-fe7c-589e-8259-30060bf20500 PFS b 0.0 1 1 0 1
[...]
raw PFS/raw/all 570092eb-f571-5631-8d20-11acbeabc640 PFS r 0.0 3 1 1 26
raw PFS/raw/all f8e3ae71-2cdf-5e55-bc42-4a4fb913770c PFS r 0.0 4 1 1 27
Datasets can be accessed from Python using the butler API:
from lsst.daf.butler import Butler
butler = Butler.from_config($DATASTORE, collections="PFS/raw/sps")
raw = butler.get("raw", instrument="PFS", visit=12, arm="r", spectrograph=1)
rawImage = raw.getImage()
The raw data returned from the butler is of type PfsRaw, which provides a common interface for both CCD and NIR detectors.
You can use butler.get("raw.exposure", ...) to get the exposure from the raw data directly.
The Default Collection
You can create a CHAIN collection, PFS/defaults, to combine all of the collections created earlier:
butler collection-chain $DATASTORE PFS/defaults PFS/raw/pfsConfig PFS/raw/sps PFS/calib
This is the collection that users will probably use for most pipeline tasks.