Analysis guides
Preparing your data
Most of the reasons an analysis fails are settled before the data ever arrives.
File formats
| Data | Format accepted | Notes |
|---|---|---|
| Raw sequencing | FASTQ (gzip fine) | see below for the filename convention |
| Expression matrix | CSV / TSV | first column gene names, remaining columns samples |
| Sample metadata | CSV | first column sample ID — must match the matrix column names exactly |
Where the raw data goes
AURORA does not accept FASTQ from the screen. The raw data has to be in your own account folder under the AURORA data folder on the server, with a project folder created inside it. Selecting the account folder itself as input is blocked.
Folder names start with a letter or digit and use only . _ -. A path set up as a symlink
will not open — it has to be the real path.
FASTQ filenames
The rules for matching pairs are fixed. Read 1 has to be one of the following, and read 2 has to match it in the same form.
.1.fastq R1.fastq _R1.fastq _1.fq .1_val_1.fq .R1_val_1.fq
A trailing .gz is fine.
10x single cell follows a different rule. It uses exactly the form Cell Ranger requires:
<sample>_S1_L001_R1_001.fastq.gz
Pointing at the whole ..._10X_RawData_Outs folder the vendor gave you is usually the right move.
Even when it is split across several flow cells, it is merged and processed as one.
There is a reason the rules are strict. Match them loosely and biologically different samples such as
NT_cDC1_1andNT_cDC1_2get merged as read 1 and read 2 of a single sample, and an averaged value comes out without an error. Better to be stopped than to be quietly wrong.
The sample metadata table
AURORA builds a table from the samples it detected. These are the fields you fill in.
| Field | Meaning |
|---|---|
sample_id |
Sample name. It becomes a folder name at every step |
fastq_sample |
The value from the scan — you cannot edit it |
patient_id |
Patient or subject identifier |
group |
Comparison group. Record the group used for between-group comparisons |
tissue_type |
Tissue |
batch |
Batch |
Because sample_id becomes a folder name, use only letters, digits and . _ -. Blanks or
duplicates block the run.
Somatic variants are paired by name
Normal and tumour are told apart not by a separate field but by the characters in the sample
name. The defaults are ASN for normal and ASC for tumour, and you can change them on screen.
If the name does not contain those characters, no pair is made and somatic variant calling does
not run.
Where people get stuck
Sample names do not line up. If a matrix column name and a metadata ID differ by even one
character, that sample is dropped silently. Check spaces, capitalisation, and - versus _.
There is only one group. Differential expression needs at least two groups to compare.
There are no replicates. One sample per group leaves nothing to do statistics on. Two is the minimum; three or more is better.
Exome, but no capture BED. Mutect2 searches the whole genome and the CNVkit reference gets noisy. If you have it, supply it.
CNVkit is on but there are few normal samples. CNVkit pools normal samples to build a reference. Five or more is recommended. With none, the run is blocked.
Check before you upload
- Does the file open at all (is the archive intact)
- Is the sample count the number you expected
- Are all the pairs matched — samples with no mate pass by as a warning only
- How much is missing — above half, the sample is usually better left out
Last updated ·
Docs