FastQ format
For more information on the FastQ format, see the MAQ FastQ documentation.
FastQ is the standard format for storing the raw reads that come off a sequencer. It extends the FastA format by pairing each base with a quality score that describes how confident the sequencer was in that basecall.
Format
Each read in a FastQ file is described by four lines:
A line beginning with
@that holds the sequence identifier and an optional description.The raw sequence (the bases).
A line beginning with
+, optionally followed by the same identifier again.The quality scores, encoded as ASCII characters, one per base in line 2.
Example
A single FastQ read looks like this:
@HWI-ST911:111:C0N4WACXX:5:1101:2249:2216 1:Y:18:TTAGGC
TTAGGCAGGACAGCTCAGGGCATGAAGTTGTTAATTCAGGACAGGGCATGT
+
CCCFFFFFHHHHHJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJHIJJJ
The identifier line packs a lot of information. Breaking the first example apart:
HWI-ST911— the unique instrument name111— the run idC0N4WACXX— the flowcell id5— the flowcell lane1101— the tile number within the lane2249— the x-coordinate of the cluster within the tile2216— the y-coordinate of the cluster within the tile1— the member of a pair (1 or 2; paired-end reads only)Y— whether the read failed the filter (Y = filtered out, N = passed)18— control bits (0 when none are on)TTAGGC— the index (barcode) sequence
What software use FastQ?
FastQ is the entry point for almost every downstream tool. A few examples:
How are these files generated?
FastQ files are produced by the sequencer’s basecalling software, which converts the raw signal measured for each cluster into a sequence of bases and their associated quality scores.