Docs

Services

AURORA

Automated Unified Reproducible Omics Reporting & Analysis

An NGS pipeline that takes raw FASTQ already on the server through QC, alignment, variant calling and expression quantification, and produces a report.

Open AURORA

What it does

AURORA (Automated Unified Reproducible Omics Reporting & Analysis) takes raw sequencing data and processes it so that the result does not change when the person does. It records every command it ran, without exception, and reads tool versions back out of the outputs, so months later you can retrace how those numbers were produced.

What to know before you start

The raw data must already be on the server. There is no FASTQ upload on this screen. You put the data in your own account folder under the designated data folder on the server, then pick that folder on screen. The only thing you can upload is a sample metadata CSV.

The reference genome is fixed to hg38. No other build and no other species can be selected right now.

Reference files and tools have to be installed on the server beforehand. Nothing is downloaded while a job runs. What is present and what is missing is shown as-is under 환경 점검 (Environment check).

Not everything you can see on screen actually runs. If a tool is not installed you can still select it, but the run will be blocked. After choosing a preset, always check 환경 점검 (Environment check) on the right. Which program each step uses and what is ready right now is set out in Software availability.

How you use it

The screen has four sections in the left sidebar — 새 분석 (New analysis) · 작업 현황 (Job status) · 결과 · 리포트 (Results and reports) · 관리자 (Admin) (admin only appears if you have the permission).

  1. Source folder — create a folder with 프로젝트 만들기 (Create project) or pick an existing project. Drill down to where the raw data is. Samples are detected and shown as soon as you select it
  2. Sample metadata — the detected samples appear as a table. Click a cell to edit it in place. You can also download it with CSV 내려받기 (Download CSV), fill it in, and send it back with CSV 올리기 (Upload CSV). If you fill in the comparison group (group), you can start a between-group comparison later on
  3. Pipeline — choose the data type and a preset. Whatever the preset fixes is locked; you choose the rest step by step
  4. Report details — client, institution, sequencer and so on. Blank fields are left out of the report
  5. Run plan and environment check — the right-hand column shows what will run in what order and whether the required tools are installed, in real time
  6. Start the analysis — if there is even one error the button will not respond
  7. Job status — refreshes every 5 seconds. The analysis keeps running if you close the window. You can open a row, cancel a run, or delete it
  8. Results and reports — seven tabs show the output

The seven result tabs

Tab What it shows
진행 상황 (Progress) Progress and per-step status
결과 파일 (Result files) Outputs per step, file listing, 전체 내려받기 (zip) (Download all as zip)
실행 명령어 (Commands run) Every shell command the pipeline actually executed
실행 로그 (Run log) Standard output and standard error
품질 검사 (QC) Four figures (sequencing depth, GC content, trimming removal rate, mapping rate), FastQC reports, per-module verdicts
리포트 (Report) Edit the details, rebuild, and download as HTML or PDF

What is implemented

Data types accepted — WGS (whole genome) · WES (whole exome) · BarSEQ (cfDNA barcodes) · RNA-seq (transcriptome) · miRNA-seq (small RNA) · scRNA-seq (10x single cell).

Pipeline presets — GDC standard, GATK standard, somatic variants (BWA-MEM2 + Mutect2), BarSEQ (cfDNA), Tuxedo, STAR + featureCounts, miRSEQ, Cell Ranger. Only the ones that fit the data type are shown — for single cell there is only Cell Ranger. You can also pick no preset and specify everything yourself.

Tools you can choose per step

Step Options
Quality check FastQC
Trimming Trim Galore · Cutadapt · off
Alignment BWA (mem · mem2 · aln) · STAR · TopHat2 · Bowtie2 · off
Duplicate removal Picard MarkDuplicates · barcode-based · off
Realignment GATK IndelRealigner
Variant calling GATK HaplotypeCaller · VarScan2 · MuTect2 · Mutect2 (GATK4 somatic) · MuSE · SomaticSniper
Variant annotation ANNOVAR · VEP
Expression quantification HTSeq · Cufflinks · featureCounts
Copy number variation Sequenza · CNVkit
Single cell Cell Ranger count

Command preview mode — records the commands that would run without actually running them. Use it to check your settings, or to get commands to run on another server.

Reproducibility record — every command run, tool versions (read back from the output files, and asked of the installed tool directly when they are not there) and 24 run settings go into the report appendix.

Reports — HTML and PDF. Eight sections, and they are produced even when the pipeline fails — the point is to show you how far it got.

Admin features — force a re-scan of the environment check, override tool and reference paths.

What it cannot do

You cannot upload FASTQ. Moving data onto the server has to happen outside this screen.

You can only work under your own account folder. Selecting the user folder itself as input is also blocked — you must pick a project folder beneath it.

Symlinked paths do not work. The path check resolves to the real path before comparing, so a folder set up as a link will not open.

There is no queue. Press 분석 시작 (Start analysis) and it runs immediately. Launch several and that many run at once — you have to manage resources yourself with 동시 실행 시료 (Concurrent samples) and the thread count.

There is nothing but hg38. Neither the reference build nor another species can be changed on screen.

Do not download large files through the browser. Downloads load the whole file into server memory before sending it, so for a large BAM you are better off taking it from the filesystem directly.

Switching to English leaves most of it in Korean. Table headers and cell values, buttons and notifications are fixed in Korean. The language switch only reaches the messages the server generates and some option names.

Some settings are not on screen. The Cufflinks reference GTF, the variant depth cutoff, the Cell Ranger core and memory settings and others run on defaults only. Changing them means an administrator has to put them in through a path override.

Cancelling a run is not complete. Only the runner and its direct children get the termination signal. Grandchild processes already running, such as STAR or BWA, can survive.

Rebuilding a report does not tell you the outcome. You only get a notification that the build started, and it stays quiet if it fails. Reopen the tab and check whether the file changed yourself.

Where people get stuck

It cannot find the FASTQ files. Check the filename convention. Read 1 has to look like .1.fastq · R1.fastq · _R1.fastq · _1.fq, and read 2 has to match it. 10x single cell requires the <sample>_S1_L001_R1_001.fastq.gz form — check that you pointed at the ..._10X_RawData_Outs folder the vendor gave you, as it is.

It says some samples have no mate. That is a warning, not an error. If the data is not single-end, check the filenames.

Trimming is off and alignment is blocked. BWA, TopHat2 and Bowtie2 read their input from the trimming output folder. Either turn trimming on or use STAR. STAR is the only exception.

You cannot turn off both quality check and trimming. The pipeline cannot build the input FASTQ list. At least one has to be on.

Somatic variants, but no pairs are found. Normal and tumour are told apart by a tag in the sample name. The defaults are ASN for normal and ASC for tumour. If the name does not contain those characters, no pair is made.

Exome without a BED file. Mutect2 searches the whole genome and the CNVkit reference gets noisy. Supply the capture BED.

Editing a sample ID causes an error. The sample ID becomes a folder name at every step. Use only letters, digits and . _ -. Duplicates and blanks block the run too.

A job is left as orphaned. The status file says it is running but there is no process. This happens when the server was restarted — delete it and start again.

The first request is slow. Reading the configuration takes about 7 seconds. Once per worker; after that it comes from cache. It is not downloading a reference genome.

Reading the output

Read QC summary → alignment rate → expression and variants, in that order. Which numbers you can trust is covered in Reading your results.

Software availability

AURORA does not do the computing itself. At every step it calls out to an outside program — BWA and STAR do the alignment, GATK does the variant calling. So if that program is not installed on the server, that step does not run.

환경 점검 (Environment check) on the AURORA screen tells you that before you run. This section is about how to read that screen, and what is ready right now.

The original paper for each program is collected in the References below.

Where to look on screen

Where What you see
The right-hand column of 새 분석 (New analysis) Only what the pipeline you picked actually needs
The 관리자 (Admin) tab All 23 tools and all 18 reference files. Administrators only

The 새 분석 side is the one that matters. If even one of the required entries reads 없음 (missing), the 분석 시작 (Start analysis) button will not respond. It stops you before the start rather than halfway through — finding out the next tool is missing after hours of alignment is far worse.

What each step uses

The statuses below are as of 2026-08-24. The source of truth is always 환경 점검 on screen, so check there right before you run. When a tool is newly installed, that screen changes before this table does.

Step Program Status
Quality check FastQC Available
Trimming Trim Galore Available
Trimming Cutadapt Not available
Alignment BWA (mem · mem2 · aln) Not available — SAMtools is needed alongside it
Alignment STAR Available
Alignment TopHat2 Available
Alignment Bowtie2 Not available
Shared SAMtools Not available
Duplicate removal Picard MarkDuplicates Available
Realignment GATK 3 IndelRealigner Available
Variant calling GATK 4 HaplotypeCaller Available
Variant calling GATK MuTect2 · GATK4 Mutect2 Available
Variant calling VarScan2 Not available — SAMtools is needed alongside it
Variant calling MuSE Not available
Variant calling SomaticSniper Not available
Variant annotation ANNOVAR Available
Variant annotation VEP Available
Expression quantification Cufflinks Available
Expression quantification HTSeq Not available
Expression quantification featureCounts Not available
Copy number variation CNVkit Available
Copy number variation Sequenza Not available — SAMtools is needed alongside it
Single cell Cell Ranger Not available
Building R objects SEQprocess (R) Available — it does not use an outside program

Note that SAMtools runs across several rows. Alignment (BWA), VarScan2 and Sequenza all call it alongside their own work. So one missing tool blocks three steps at once.

What can run end to end right now

Of the nine presets, only one runs right now: Tuxedo (RNA-seq). Trim Galore → TopHat2 → Cufflinks are all in place.

Preset Right now What blocks it
Tuxedo (RNA-seq) Runs —
GDC standard (RNA-seq) Blocked HTSeq
STAR + featureCounts Blocked featureCounts
GDC standard (WGS/WES) Blocked SAMtools, VarScan2
GATK standard Blocked SAMtools (BWA alignment)
Somatic variants (BWA-MEM2 + Mutect2) Blocked SAMtools
BarSEQ (cfDNA) Blocked SAMtools
miRSEQ Blocked Cutadapt
Cell Ranger (10x) Blocked Cell Ranger

Put together yourself, STAR alignment + Cufflinks quantification works as well. Running only the quality check works too.

명령어 미리보기 (Command preview) is not subject to this. It is a mode that records the commands without actually running them, so it goes all the way through even with a tool missing. Use it to check your settings, or to take away commands to run on another server.

Reference files are checked too

Not just programs — reference files come up in the same table: the reference genome FASTA, the alignment indexes, VCFs such as dbSNP, COSMIC and gnomAD, the ANNOVAR and VEP databases, 18 of them in all.

This is what has been confirmed so far.

Reference Status
Reference genome FASTA (hg38) Ready
BWA index Ready
STAR index Ready
ANNOVAR database Ready
UCSC refFlat (CNVkit gene annotation) Missing

The exome capture BED is empty by default. Its absence does not block a run, but you have to supply it when you use Mutect2 or CNVkit on exome data — without it Mutect2 searches the whole genome and the CNVkit reference gets noisy. For the details see Preparing your data.

What to do when you are blocked

Installing is not something you can do. Programs and reference files have to be on the analysis server, and that server is the administrator's to touch.

  1. In 환경 점검, note exactly which names read 없음
  2. Pass those names to an administrator
  3. If you are in a hurry, see whether another tool gets you around it — there is usually more than one option for the same job. If expression quantification is blocked, go to Cufflinks; if variant annotation is blocked, switch from ANNOVAR to VEP
  4. If you do not need the results this minute, settle the settings with 명령어 미리보기 and run the same settings again once the tools are in place

For administrators: the actual install paths and how to override them are in the operations wiki. A program can be installed and still read 없음 — the place the configuration points at and the place it is actually installed are different, and the fix then is not a fresh install but 경로 덮어쓰기 (Path override) in the 관리자 tab.

References

35 items

AURORA computes nothing itself — every step shells out to an external program. Below are the original papers for those programs and for the reference data.

Cite only the steps that actually ran. The result screen keeps every command the run executed, and the tool versions are read back into the report. For what is installed on the server, see Software availability.

Quality control and trimming

  • FastQC — Andrews S. FastQC: a quality control tool for high throughput sequence data. Babraham Bioinformatics. bioinformatics.babraham.ac.uk
  • Trim Galore — Krueger F. Trim Galore: a wrapper around Cutadapt and FastQC. Babraham Bioinformatics. bioinformatics.babraham.ac.uk
  • Cutadapt — Martin M. Cutadapt removes adapter sequences from high-throughput sequencing reads. EMBnet J. 2011;17(1):10-12. doi:10.14806/ej.17.1.200

Alignment

  • BWA (aln · bwasw) — Li H, Durbin R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics. 2009;25(14):1754-1760. doi:10.1093/bioinformatics/btp324
  • BWA-MEM — Li H. Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM. arXiv:1303.3997. 2013. arxiv.org/abs/1303.3997
  • BWA-MEM2 — Vasimuddin M, Misra S, Li H, Aluru S. Efficient architecture-aware acceleration of BWA-MEM for multicore systems. IEEE International Parallel and Distributed Processing Symposium (IPDPS). 2019. github.com/bwa-mem2
  • Bowtie2 — Langmead B, Salzberg SL. Fast gapped-read alignment with Bowtie 2. Nat Methods. 2012;9(4):357-359. doi:10.1038/nmeth.1923
  • STAR — Dobin A, Davis CA, Schlesinger F, et al. STAR: ultrafast universal RNA-seq aligner. Bioinformatics. 2013;29(1):15-21. doi:10.1093/bioinformatics/bts635
  • TopHat2 — Kim D, Pertea G, Trapnell C, et al. TopHat2: accurate alignment of transcriptomes in the presence of insertions, deletions and gene fusions. Genome Biol. 2013;14(4):R36. doi:10.1186/gb-2013-14-4-r36
  • SAMtools — Li H, Handsaker B, Wysoker A, et al. The Sequence Alignment/Map format and SAMtools. Bioinformatics. 2009;25(16):2078-2079. doi:10.1093/bioinformatics/btp352 · Danecek P, Bonfield JK, Liddle J, et al. Twelve years of SAMtools and BCFtools. Gigascience. 2021;10(2):giab008. doi:10.1093/gigascience/giab008

Duplicate marking and realignment

  • Picard — Broad Institute. Picard toolkit. broadinstitute.github.io/picard
  • GATK — McKenna A, Hanna M, Banks E, et al. The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data. Genome Res. 2010;20(9):1297-1303. doi:10.1101/gr.107524.110

Variant calling

  • GATK4 HaplotypeCaller — Poplin R, Ruano-Rubio V, DePristo MA, et al. Scaling accurate genetic variant discovery to tens of thousands of samples. bioRxiv. 2018. doi:10.1101/201178
  • MuTect — Cibulskis K, Lawrence MS, Carter SL, et al. Sensitive detection of somatic point mutations in impure and heterogeneous cancer samples. Nat Biotechnol. 2013;31(3):213-219. doi:10.1038/nbt.2514
  • Mutect2 — Benjamin D, Sato T, Cibulskis K, Getz G, Stewart C, Lichtenstein L. Calling somatic SNVs and indels with Mutect2. bioRxiv. 2019. doi:10.1101/861054
  • VarScan2 — Koboldt DC, Zhang Q, Larson DE, et al. VarScan 2: somatic mutation and copy number alteration discovery in cancer by exome sequencing. Genome Res. 2012;22(3):568-576. doi:10.1101/gr.129684.111
  • MuSE — Fan Y, Xi L, Hughes DS, et al. MuSE: accounting for tumor heterogeneity using a sample-specific error model improves sensitivity and specificity in mutation calling from sequencing data. Genome Biol. 2016;17(1):178. doi:10.1186/s13059-016-1029-6
  • SomaticSniper — Larson DE, Harris CC, Chen K, et al. SomaticSniper: identification of somatic point mutations in whole genome sequencing data. Bioinformatics. 2012;28(3):311-317. doi:10.1093/bioinformatics/btr665

Variant annotation

  • ANNOVAR — Wang K, Li M, Hakonarson H. ANNOVAR: functional annotation of genetic variants from high-throughput sequencing data. Nucleic Acids Res. 2010;38(16):e164. doi:10.1093/nar/gkq603
  • VEP — McLaren W, Gil L, Hunt SE, et al. The Ensembl Variant Effect Predictor. Genome Biol. 2016;17(1):122. doi:10.1186/s13059-016-0974-4

Expression quantification and differential expression

  • Cufflinks — Trapnell C, Williams BA, Pertea G, et al. Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation. Nat Biotechnol. 2010;28(5):511-515. doi:10.1038/nbt.1621
  • HTSeq — Anders S, Pyl PT, Huber W. HTSeq — a Python framework to work with high-throughput sequencing data. Bioinformatics. 2015;31(2):166-169. doi:10.1093/bioinformatics/btu638
  • featureCounts — Liao Y, Smyth GK, Shi W. featureCounts: an efficient general purpose program for assigning sequence reads to genomic features. Bioinformatics. 2014;30(7):923-930. doi:10.1093/bioinformatics/btt656
  • DESeq2 — Love MI, Huber W, Anders S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol. 2014;15(12):550. doi:10.1186/s13059-014-0550-8

Copy number variation

  • CNVkit — Talevich E, Shain AH, Botton T, Bastian BC. CNVkit: genome-wide copy number detection and visualization from targeted DNA sequencing. PLoS Comput Biol. 2016;12(4):e1004873. doi:10.1371/journal.pcbi.1004873
  • Sequenza — Favero F, Joshi T, Marquard AM, et al. Sequenza: allele-specific copy number and mutation profiles from tumor sequencing data. Ann Oncol. 2015;26(1):64-70. doi:10.1093/annonc/mdu479

Single cell

  • Cell Ranger — 10x Genomics. Cell Ranger. 10xgenomics.com · Zheng GXY, Terry JM, Belgrader P, et al. Massively parallel digital transcriptional profiling of single cells. Nat Commun. 2017;8:14049. doi:10.1038/ncomms14049

Pipeline framework and report

  • SEQprocess — Joo T, Choi JH, Lee JH, et al. SEQprocess: a modularized and customizable pipeline framework for NGS processing in R package. BMC Bioinformatics. 2019;20(1):90. doi:10.1186/s12859-019-2676-x · github.com/omicsCore/SEQprocess
  • R — R Core Team. R: a language and environment for statistical computing. R Foundation for Statistical Computing. r-project.org
  • Shiny — Chang W, Cheng J, Allaire JJ, et al. shiny: web application framework for R. shiny.posit.co
  • R Markdown — Xie Y, Allaire JJ, Grolemund G. R Markdown: the definitive guide. Chapman & Hall/CRC; 2018. rmarkdown.rstudio.com

Reference genome and databases

  • GRCh38 (hg38) — Schneider VA, Graves-Lindsay T, Howe K, et al. Evaluation of GRCh38 and de novo haploid genome assemblies demonstrates the enduring quality of the reference assembly. Genome Res. 2017;27(5):849-864. doi:10.1101/gr.213611.116
  • dbSNP — Sherry ST, Ward MH, Kholodov M, et al. dbSNP: the NCBI database of genetic variation. Nucleic Acids Res. 2001;29(1):308-311. doi:10.1093/nar/29.1.308
  • COSMIC — Sondka Z, Dhir NB, Carvalho-Silva D, et al. COSMIC: a curated database of somatic variants and clinical data for cancer. Nucleic Acids Res. 2024;52(D1):D1210-D1217. doi:10.1093/nar/gkad986
  • gnomAD — Karczewski KJ, Francioli LC, Tiao G, et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature. 2020;581(7809):434-443. doi:10.1038/s41586-020-2308-7

Last updated ·