Thursday, 2 February 2017

Understanding your @10Xgenomics cell ranger reports

The Cell Ranger analysis provided by 10X is an excellent start to understanding what might be going on in the single cells you just sequenced. It allows some basic QC and this can help determine how well your experiment is working. There is a high degree of variability in the number of cells captured and capture efficiency, but right now we cannot easily see if this is down to the sample (most likely) or the technology.



Some of the metrics are easy to interpret e.g. the ‘Estimated Number of Cells’ (how many single cells were captured) – the more the merrier! Others need to be compared across runs to determine what the “correct” parameters for an experiment might be e.g. the current 10X recommendation for ‘Mean Reads per Cell’ is 50,000, but you may find that more, or fewer, reads are required for your samples. You can use the other metrics such as ‘Median Genes per Cell’ or ‘Sequencing Saturation’ to help determine when more or less sequencing depth are required.

The most important metrics: 10X help by making the most important stuff big. You should already have an idea of the number of cells you expected to capture (because you carefully counted your cells before starting didn't you), hopeful the  ‘Estimated Number of Cells' matches what you were aiming for. Ideally this would be the same across your project, but is likely to be quite variable if the cell types are very different.

The ‘Mean Reads per Cell’ and ‘Sequencing Saturation’ both tell you whether you've over-sequenced. Our recommendation is to run a single lane on hiSeq 400 first and to use these numbers to determine if more sequencing is worth it or not. Diving in for a lane per sample might turn out to be expensive mistake (as it was in the example above).

The ‘Median Genes per Cell’ equals is likely to become a key metric for users. We've become used detecting 10,000-15,000 genes in microarray and RNA-Seq experiments on bulk tissue. What the figure is for single-cell remains to be seen. However it is likely to be quite cell specific, and is also likely to increase as methods capture more of the transcripts.


The ‘Sequencing’ table metrics explained:
  • ‘Number of Reads’ equals the total number of single-end reads that were sequenced.
  • ‘Valid Barcodes’ equals the fraction of reads with barcodes that match the whitelist.
  • ‘Reads Mapped Confidently to Transcriptome’ equals the fraction of reads that mapped to a unique gene in the transcriptome with a high mapping quality score as reported by the aligner.
  • ‘Reads Mapped Confidently to Exonic/Intronic/Intergenic Regions’ equals the Fraction of reads that mapped to the exonic/intronic/intergenic regions of the genome with a high mapping quality score as reported by the aligner.
  • ‘Sequencing Saturation’ equals the fraction of reads originating from an already-observed UMI. This is a function of library complexity and sequencing depth. More specifically, this is the fraction of confidently mapped, valid cell-barcode, valid UMI reads that had a non-unique (cell-barcode, UMI, gene). This metric was called "cDNA PCR Duplication" in versions of Cell Ranger prior to 1.2.
  • ‘Q30 Bases in Barcode/Sample Index/UMI Read equals the fraction of bases with Q-score at least 30 in the cell barcode/sample index/Unique molecular identifier sequences.
  • ‘Q30 Bases in RNA Read’ equals the fraction of bases with Q-score at least 30 in the RNA read sequences. This is Illumina R1 for the Single Cell 3' v1 chemistry and Illumina R2 for the Single Cell 3' v2 chemistry.
  • ‘Estimated Number of Cells' equals the The total number of barcodes associated with cell-containing partitions, estimated from the barcode count distribution.
  • ‘Fraction Reads in Cells' equals the The fraction of barcoded, confidently mapped reads with cell-associated barcodes.
  • ‘Mean Reads per Cell’ equals the total number of sequenced reads divided by the number of barcodes associated with cell-containing partitions.
  • ‘Median Genes per Cell’ equals the median number of genes detected per cell-associated barcode. Detection is defined as the presence of at least 1 UMI count.
  • ‘Total Genes Detected’ equals the number of genes with at least one count in any cell.
  • ‘Median UMI Counts per Cell’ equals the median number of UMI counts per cell-associated barcode.

It will help to look at these numbers over time and across projects. Right now the data about the sample is limited, but collecting more sample/experiment metadata is likely help determine whether an experiment has worked or not. Right now it is difficult for us to give advice as your experiment may be the first time we've eve3r run that type of cell!

Monday, 24 October 2016

How do I submit my index information into Lablink?

When accepting sequencing submissions in the Genomics Core, there may be instances where we have to contact you if there is an error with your submission form. The most common problems relate to index information. We have put together some instructions here that we hope should make things easier and help us to get started on your sequencing as soon as we can.

Please only follow these instructions if:
·        The index sequences you have used are visible in the index sequences tab of the submission form
·        There are fewer than 384 samples within your pool.
If the points above are not true, please see section, unspecified index further on in this blog.

1.      Completing the sample/reagent label field

1a. Navigate to the index sequences tab of the sample submission form.

1b. Search for your index sequences

1c. Copy the index name from column C of the index sequences tab, e.g A001-A005 to the column Sample/Reagent Label, of the submission form tab.

Figure 1-index sequences tab of the sample submission form

Figure2-submission form

      2. Completing the UDF/Index type field

2a. Select the correct UDF/Index type from the drop down menu on the submission form tab.
IMPORTANT –please make sure the Index type field matches column B of the index sequences tab. This ensures that your library goes through our acceptance step. Please see the two following examples. 

Example1- I am submitting a Truseq LT library consisting of 5 samples and used indexes A001-A005.
The sample/reagent label on the submission form should read A001-A005.
The UDF/index type should read Truseq LT.

 Figure 3 - The index type field next to these indexes is Truseq LT so this is what should be entered into the UDF/Index type field.



Figure 4- Submission form

In most cases, the Index type will match to the indexes you have used as expected. However there are now multiple kits available which share the same indexes.
Because of this, there may be some cases where the index sequences you select will have a different Index type to the library you have made. (see example 2 below) This may affect you if you are submitting for Nextera XT or Nextera.

Example 2- I used indexes N701-N501, N702-N501, N703-N501, N704-N501. I prepared the libraries using a Nextera XT library prep kit.

The sample/reagent label on the submission form should read N701-N501, N702-N501, N703-N501, N704-N501. The UDF/index type should read Nextera and not Nextera XT. This is because the UDF/Index type needs to match column B on the index sequences tab.


Figure 5 - The index type field next to these indexes is Nextera so this is what should be entered into the UDF/Index type field

 
3. Unspecified Index

If your pool has index sequences not present in the index sequences tab OR if you have a pool which is made up of more than 384 samples, you will need to submit as unspecified index.
In the submission form:
·        Sample/reagent label should read – unspecified
·        UDF/Index type should read –Unspecified (other)
You should submit your pool as one row on the form. Libraries submitted as unspecified index cannot be demultiplexed by the Genomics Core but we do have a demultiplexing guide on lablink which should give some useful information.

Important-since we have no index sequence information, please write in the comments section of the form the index lengths for Index 1 and Index 2. Without this information your sequencing may be delayed whilst we contact you to check these parameters.
Once you have submitted your libraries, the Genomics Core would like to start working on your sequencing as soon as we can.

If the incorrect index type has been selected, we will need to delete your submission and we would ask you to submit again after making changes to your sample sheet and following the instructions above. Of course whilst this guide should be used to help you, we are always here to discuss this with you in person if you have any questions. Alternatively you can contact us on our helpdesk: genomics-helpdesk@cruk.cam.ac.uk
 



Sunday, 9 October 2016

Recent papers that the Genomics Core has helped with

I like to highlight some of the really interesting work we've been involved with, or that has come out of the Institute from time to time, and I recently updated our lab home page with links to a couple of papers.  i thought I'd take the opportunity to write about them in a bit more detail here. Many of you will already know I run the Genomics Core facility at CRUKs Cambridge Institute. We do a lot of Illumina sequencing! The lab works on a huge number of projects for the research groups here in the Institute, and also across many groups in Cambridge via a long-running sequencing collaboration. We do do some R&D work in my lab, but >90% of our efforts are working with, or for, other research groups.
Highlights from the last years genomics research include work from the Caldas group who have completed three project over the lat year I've included here; 1) profiling of almost 2500 Breast Cancer patients for mutational analysis of 173 genes using a targeted pull-down (Pereira et al Nature Communications 2016); 2) cancer exomes from Murtaza et al,; 3) PDXs from Bruna et al.; and the Balasubramanian group who have shown that it is possible to capture and sequence double-strand DNA breaks (DSBs) in situ and directly map these at single-nucleotide resolution, enabling the study of DSB origin (Lensing et al. Nature Methods 2016). The rapid speed and unbiased nature of the genome-wide experiments being performed in the Institute, and often prepped and sequenced in the Genomics core continue to increase our understanding cancer biology.


Friday, 15 July 2016

Why is my HiSeq 2500 sequencing taking longer than usual

With the introduction of the HiSeq 4000 we're able to sequence faster and cheaper than ever before. But as we're transitioning the larger projects over to HiSeq 4000 a side-effect is fewer and fewer samples to run on HiSeq 2500; and as we're waiting for samples to fill the 8 lane flowcell that means longer wait times for you. We thought this post might help you determine if you still need to use HiSeq 2500, or if you can migrate over to HiSeq 4000. Most sequencing is taking under 2 weeks, but some people are now waiting up to one month for 2500 data.



Running a big RNA-seq project is easy(ish)

Last year we completed our largest ever RNA-seq project: 528 samples of TruSeq mRNA, 60 lanes of HiSeq 2500 SE50, 13 billion reads - and all in 16 weeks. Being able to do such a large project in such a short time and get high quality data from nearly all samples really demonstrates the robustness of RNA-seq. If you're thinking that a project larger than 96 samples might be too much to consider, then come and talk to us (and Bioinformatics) at a Tuesday afternoon experimental design meeting - and we'll convince you it can be a pretty smooth process.



We've been using Illumina's TruSeq mRNA-seq automated on our Agilent Bravo robot and the sequencing was done on HiSeq 2500, although we're currently  moving to HiSeq 4000.
  • 528 samples processed on six-plates of RNA-seq
  • QC lanes sequenced and analysed
  • 60 lanes of SE50bp sequencing in total, 10 lanes per plate
  • 12,918,018,345 PF reads for this project (215M reads per lane on average)
  • 24M reads per sample on average
  • 16 weeks from start to finish
This has been a large and complex project where we had lots of discussions along the way. I think that everyone involved has contributed to the success so far: the research group who asked us to do the project, my lab, and also our Bioinformatics Core. The ability to discuss the experiment at different stages, and to focus on QC issues as they arise really makes using the Cores a great place to do your projects.

Sunday, 7 February 2016

Our first paper on the bioRxiv

I just uploaded our paper, which has also been submitted to BioTechniques, onto the bioRxiv preprint server. The work we present comes from an idea I had shortly after first using Agilent's BioAnalyser in 2000. I was blown away by this piece of technology that has become the de facto standard for RNA QC, and has also pretty much replaced gel electrophoresis for DNA fragment analysis in NGS applications. When launched in 1999, it was the only microfulidics instrument for biology applications. The idea was a simple one: can bioanalyser chips be swapped between assays?

Friday, 20 November 2015

Following us on Twitter

The Genomics Core now has two Twitter accounts, you can follow me @CIgenomics (James Hadfield, Head of Genomics) and hear about things I think are interesting, but which you might not necessarily be interested in; and/or you can follow our sequencing queue @CRUKgenomecore which puts out live Tweets directly from the sequencing LIMS.



How does the LIMS Tweet: Some clever work by Rich in Bioinformatics has allowed us to pull out data directly from Genologics Clarity LIMs queue using a script run every 24 hours, and the Twitter API then allows that script to post messages on our behalf. Because of this the Tweets about our queue should happen every day and without manual intervention. Hopefully you'll be able to rely on these to give you a reasonable idea of how long you might have to wait for your sequencing results. Of course we can't predict what will happen with your particular sample so please treat the Tweet as a guide.

Tweets explained: The Tweets have a format that we hope is pretty intuitive, but we've described what all the bits of information mean below...


Thanks especially to Rich Bowers in the Bioinformatics core for pulling all of this together from a vaguely described idea by me.

Friday, 4 September 2015

Improving DNA and RNA quant with plate based fluorimetry

We quantify NGS libraries all the time and qPCR works brilliantly, but nucleic acids need to be handled differently. We don't actually run that much quantifiaction on DNA and RNA as most of our users have already done this; we asked them to do it so we could more efficiently run larger batches of library prep to keep costs down and turnaround times as short as possible. Over the last few years we've been running the Nextera exome preps and DNA quant has become more important than ever before, in fact we started running a secondary quant just to be certain about DNA concentration.

Most of the time DNA and RNA quant works well and we've favoured the fluorescent Qubit assay recommended by Illumina in their protocols. A nanodrop or plate reading spec at 260:280nM measures total nucleic acid and is confounded by ssDNA, RNA, and oligos so can give inaccurate results. We run the Qubit dsDNA BR Assay from Molecular Probes on the PHERAstar fluorescent plate reader (here's their handy protocol). We have only been using 1ul of DNA (Illumina suggest 2) for each sample but we run triplicate assays to get a high-quality quantitation.

Problems with the Qubit assay: Recently some users have reported problems with the accuracy of the QuBit assay on our plate reader and the manager of our Research Instrumentation Core helped us to get to the bottom of the issues and some excellent results. The main problem turned out to be addition of DNA into the working dye solution, it was the DNA coating the outside of the tips that appeared to be making the results so flaky. Changing the protocol to add DNA to the plate first fixed it and the results are looking great.

It ca also be very important to be certain which assay you should use; BR (Broad range) or HS (High Sensitivity). If you are working with low concentration nucleic acids then the HS assay is probably the one to use. For really accurate quant we'd suggest a quick QT check first, then normalisation of samples to about twice what you need; a second triplicate and robust quant will allow you to dilute the samples to the perfect working concentration.

Here are our top tips:
  • Add DNA to the measurement plate/tubes before anything else
  • Use a repeat pipette to make sure each well gets the same/right amount of dye solution
  • Shake the tubes/plate in the dark for at least 10 minutes (quant will be inaccurate if the dye has not intercalated properly, you can check your standard curve replicates to verify if this is an issue)
  • The triplicates really are worth the effort - especially if you're doing a Nextera prep

Tuesday, 25 August 2015

When will my sequencing be done?

Will my sequencing be done before the dying of the sun,
Will wildcats once more roam the land
Will the desert still have sand
Will Norfolk be swallowed by the sea
Do I have time for a cup of tea?
Will rhinos and the manatee
Be urban legend, just like me
Oh, it's done.

Thursday, 26 March 2015

Nature reports on "careers in a core lab"

In this weeks issue of Nature a feature by Julie Gould covers what life as a core lab manager is like: Core facilities: Shared support. She interviews several core lab managers/directors from the US and Europe including me. If you've ever fancied a job in a core then I'd recommend the article.

If you have any questions about the realities of running a core and what sort of career move it might be feel free to get i touch. If you are in the CRUK-CI then you've got lots of other core managers who can give you there views as well.

Friday, 6 February 2015

Is your antibody any good

"Doesn't necessarily do what it says on the tin!" 

Probably not is the simple answer, and only if you've verified it is a more comprehensive one. The lack of reproducibility from antibody data in scientific publications is shocking, Nature published a commentary signed by over 100 researchers: Reproducibility: Standardise antibodies used in research, in which they describe the pretty poor state of antibody reproducibility. In this they cite a 2008 BioTechniques article, and a 2012 Nature commentary that discuss the state of affairs with antibodies in particular, and with reproducibility in general. In the BioTechniques paper the authors finish by saying that "for the meantime, however, the responsibility ultimately lies with the researcher or laboratory director to ensure that the antibodies used in their labs are validated for specificity and reproducibility."

Antibodies sold as being specific for a protein are oftentimes not, they can be very promiscuous in what else they bind and sometimes don't even bind the targeted protein. To make sure you are not affected by poor choices of antibodies make sure you run some validation studies before diving into your ChIP-seq experiments!

Not doing this risks wasting money (a lot according to the Nature article -see figure below). But more importantly you might waste your time, or even worse publish something that is erroneous. Hopefully you've already validated that your MCF7 cells are actually MCF7s with the BioRepository, so why not do the same with your antibody before starting your next experiment?

Figure from Bradbury and Plückthun Nature 2015.



Tuesday, 27 January 2015

Use your local support team

We have a half-day workshop on Thursday for NGS newbies, the focus of which is library prep for next-generation sequencing. We organise seminars from commercial providers of new technologies throughout the year; but this is a semi-annual event where local users get a chance to present their work, and new users get to hear about what's possible with NGS.

This year we have presentations about RNA-seq, ChIP-seq, Exome-seq, FFPE genomes, DNA methylation, targeted resequencing and a talk on the UoC 10,000 Genomes Project; and afterwards we'll wrap up with beer and pizza. These days require lots of organisation (thanks to Fatimah for organising this years event) but, for the new users especially, turn out to be well worth the effort.

Making use of your local support teams: We also make sure we keep a good relationship with our local technical support teams and run a series of commercial presentations throughout the year. This works out to be much easier to organise as they do the prep work! While we're here in the Genomics Core to help our local users, we get lots of queries from people outside the Cambridge Institute, and this is one way we've found to increase the support we can offer.

Every other month we have Illumina come in to present on a specific library prep, or talk about recent updates. Sandra (Field Application Specialist), and Carla (Marketing Technology Specialist) generally talk for 30 minutes followed by Q&A, and then spend some time with users on a one-to-one basis troubleshooting their problems.

We also try to arrange a training session once per quarter with Thermo. We've been using their ABI 7900 qPCR instruments for eight years and buy in quite a lot of their SYBR and TaqMan master-mixes. Ever since we started working with them we've run "An introduction to qPCR" course for new users. The last one was run by Emma and everyone said it was a great introductory session.

What's in it for them: Neither Illumina or Thermo would do this for free if there was nothing in it for them. They get to interact directly with potential new customers, and get feedback on how their technologies are working in the real world. Some of these conversations might end up as research collaborations. Some of the contacts might end up as new sales contracts too (I know why they are really here)!

What's in it for us: These talks have been reasonably well attended and increase the support we can offer (albeit indirectly), and the feedback from users has been almost universally positive. I'd encourage you to get in touch with your local sales or technical rep and ask if they can help you too. They might even supply doughnuts!

PS: Thanks very much to Carla and Sandra at Illumina for the seminars over the past 12 months. And to Emma for the most recent qPCR training.

PPS: If you missed the registration link to the event on Thursday, send us message via a comment below!

Sunday, 25 January 2015

How many reads do I need to sequence?

A common question we're asked is "how many reads should I use to sequence a sample?" I'm going to focus on genomes, exomes and amplicomes in this post and introduce the Lander-Waterman equation [1]. Other apps are more complex because the number is very much 'how long is a piece of string' for RNA-seq, ChIP-seq and other counting applications - it depends on the complexity of your sample and the sensitivity you'd like to get, but is also affected by the number of replicates you have.

The Lander-Waterman equation
Lander-Waterman: Almost everyone doing NGS is using this equation, even if they are not aware of it. Anyone under 27 was born after it was published (1988), but it is an equation that is good to understand if you are sequencing. Basically it allows you to estimate how many reads of a specific length you need to sequence your genome.

The general equation is C = LN/G where: C = redundancy of coverage, G is the haploid genome size, L is the sequence read length, and N is the number of sequence reads. It can be rearranged to N = CG/L allowing you to compute the number of reads to sequence a genome, exome or amplicome (amplicon-panel) to a desired coverage (this is what we typically discuss when designing experiments).

In the examples below paired-end reads of 125bp from each end of a fragment are used, but these are converted to single 250bp reads for simplicity.
  • Human genome (3Gb) 30x coverage = 360M reads.
  • Human exome (150Mb) 50x coverage = 30M reads.
  • Human amplicome (30x250bp amplicons 0.075Gb) 1000x coverage = 0.3M reads.

[1] Lander, E. S. & Waterman, S. Genomic Mapping by Fingerprinting Random Clones : A Mathematical Analysis. Genomics 239, 231–239 (1988).  
 
Eric Lander founded both the Whitehead and Broad Institutes. Michael S. Waterman is one of the founders of computational biology and gave his name to another important algorithm: Smith-Waterman alignment, he also wrote Computational Genome Analysis with our Director Simon Tavare while at the University of Southern California


Thursday, 4 December 2014

Is my NGS library any good?


We've all been there. You bought the extortionately priced kit, you ran the gels, you lovingly removed every single SPRI bead, you sweated in a lab coat for days, and finally you elute your first ever NGS libraries. The question is, how can you tell if you were wasting your time? What if your tube turns out to contain nothing but buffer? Or worse, what if it can be sequenced, but it produces nothing more than a load of expensive gobbledegook?

Never fear, if your experimental design is up to scratch, then you need only three simple quality checks to tell you if your library is a Science paper in the making, or a bit of a dud:
  1. Bioanalyzer for Library Length
  2. qPCR for Concentration
  3. Nanodrop for Chemical Contamination (optional)

1. Bioanalyzer for Size Distribution

The Agilent Bioanalyzer or Tapestation runs 1ul of your library in a microfluidics gel-like cartridge, and shows you the range of sizes in your library, as well as an estimate of library quantity.
A good Bioanalyzer trace will look different depending on the type of library you are assaying. Preferably, your library should appear as a single discrete peak approximating a bell curve. It should be larger than ~150bp, but smaller than ~700bp.
The Bioanalyzer trace is essential for detecting Illumina adapter contamination, which can be spotted is a sharp peak between 100 - 150bp. If you are a member of the CRUK Cambridge Institute, we can train you on how to run the Bioanalyzer and offer you advice on interpreting your Bioanalyzer trace.

A clean library on the Bioanalyzer: this will sequence like a dream

A problematic library on the Bioanalyzer: it will be difficult to sequence this library well.

Once you have run your library on the Bioanalyzer, use manual integration or the region table to select the entire trace and determine the average size of your library. You will need this to calculate your nanomolar concentration later. If you sequence with us, we will ask for this information at submission - it must be accurate in order for us to provide you with a high sequencing yield and quality.

Look out! Certain library prep types do not give an accurate length estimate on the Bioanalyzer due to the presence of secondary structures in the DNA (e.g. Truseq DNA PCR-free). If you're using a kit, the protocol should clearly state if this is the case - and should give you the length to use in quantification calculations.

I wouldn't recommend you use use the Bioanalyzer nmol/l concentration for multiplexing, unless you really know what you are doing - or it is explicitly recommended in your protocol or kit. After all, the Bioanalyzer nmol/l value is only accurate for quantifying certain library prep types, and it is biased by any DNA in your sample which does not contain Illumina adapters.

2. qPCR for Quantification


I like to recommend quantification of libraries by qPCR, using primers designed to target the Illumina adapters. Our NGS service currently uses the KAPA library quantification kit (LQK) for this, and we find it very reliable - but there are alternative kits out there which we haven't tested.
A high quality library should be high concentration, ideally >10nM, but also not too high concentration, ideally <100nM. 
If you find your libraries are consistently very high yield (>100nM), then it is likely that you are performing more cycles of PCR than you need; this is likely to give you unnecessarily high PCR duplicate rates in your data. Reduce your protocol 1 PCR cycle at a time until you are reliably getting 10nM - 100nM libraries. Make sure you remember to dilute your library pools to within our submission requirements, currently 10nM - 20nM.

My top tips for high quality qPCR quantification:
  • Aliquot your qPCR mastermix and your standards into single-use batches prior to first use, to avoid template contamination and the effects of repeated freeze-thaw cycles
  • Wipe down all working surfaces and pipettes with a DNA degrading cleaning agent e.g. DNA Away/DNAoff/DNAZap, before starting work
  • Make a serial dilution and take triplicate measurements, use the median concentration result
  • Check your serial dilution and your replicate measurements give highly reproducible concentration values
  • Check that your results are all comfortably within the range of your standard curve
If you use our NGS service and you choose to use the KAPA LQK, we can provide you with aliquots of the recommended DNA dilution buffer (Tris-Hcl with 0.05% tween). Also, if you are within the CRUK-CI, we offer training on how to perform real-time PCR, and you can sign out a KAPA qPCR kit from the Genomics Core to take advantage of the Institute’s bulk discount.

If you must know about the Qubit...

Other quant methods like Qubit or Bioanalyzer can be great for some library types, as long as you know what you are doing - but both will over-estimate your library concentration if you have an inefficient adapter ligation reaction. So use them with care.

Our submission guidelines are in nmol/l (nM), so if you use the Qubit you need to convert ng/ul to nM using the following equation:

x: concentration in ng/ul 
L: average library length (bp)

y: concentration in nM.

3. Nanodrop for Chemical Contamination

The Nanodrop is a quick and dirty assay for protein and chemical contaminants which interfere with sequencing - including the real killers ethanol and phenol. Test 1ul of each NGS library, preferably before you pool them for submission. I recommend you check that the 260/280 ratio is greater than 1.8, and that the 260/230 ratio is greater than 2.0. The trace should like like this:

A good Nanodrop profile

A bad Nanodrop profile. Do you see the peak at 230nm?


A library with a 260/230 ratio less than 1.8, or a 260/280 measurement less than 2.0, may cluster poorly, and therefore generate low quality data. If you're new to the library preparation process and you can spare the sample I recommend you throw this one away and start again - while paying very careful attention to each cleanup step.
Always use the recommended cleanup method, don't be tempted to swap a bead cleanup for a column, or vice versa, even if it is more convenient! That will waste your time in the long run.
If you've got a contaminant and your library is irreplaceable, consider whether your yield is sufficiently high for you to repeat the final cleanup step. If not, have a chat with your NGS provider and ask if they will try sequencing it anyway. If you sequence with us here at CRUK-CI, we will always try our best to get you sequence data - as long as you know the you run the risk of paying for a lane of data which you can't use.

Whatever happens, do NOT use the Nanodrop quantity measurement for quantifying your DNA/RNA prior to library preparation, OR your final library concentration. DON'T DO IT. This is the most easily avoidable mistake in NGS. Don't be that scientist!

I hope that is enough to get you started. As ever, if you want advice on whether your library is going to sequence well on the Illumina platform, the best place to go is your local NGS facility (if you have one), or Illumina's technical support team: techsupport@illumina.com.

Happy Sequencing!


Friday, 17 October 2014

Indexing 2: Troubleshooting a bad index balance

Indexes are one of the simplest improvements in the last five years of sequencing, with the most incredible far-reaching effects. Today I will share a complementary pair of posts tackling the problems our customers experience most frequently when submitting indexed libraries for sequencing.


Why did I get very different yields for the libraries in my pool?

We've seen this so many times. You think you have carefully quantified and pooled your libraries, and then your sequencing data comes back with a massive variation in the number of reads for each library in your pool. What a nightmare! 

Don't be fooled - there is nothing that your sequencing provider can do on the sequencer to cause a variable yield from your different indexes. An imbalance between indexes within your library pool arises during the pooling process, so an imbalanced pool indicates something has gone wrong during pooling.


Normally the problem is one of the following:


  1. Different libraries in the pool are of different lengths
  2. Quantification of the libraries prior to pooling was not accurate
  3. The process of mixing the libraries into the pool was not robust



First check #1: Are your libraries of different average size?


  • Measure the length of every library prior to pooling on the Bioanalyzer or Tapestation (or similar). 
  • Make sure you are including all of the visible peaks in your length measurement, including any adapter dimers, since they all contribute to the clustering.
  • Check that all of the libraries in your pool are a similar length to one another

Clustering efficiency is a non-linear function of length, because small fragments cluster disproportionately more efficiently than large ones. So if you mix a library of 200bp 50:50 with a library of 600bp, you will receive much more data for the short 200bp library.

As a guideline, all libraries should ideally be within +/- 50bp of one another. 


Then check #2: Was your quantification prior to pooling accurate?

If your quantification is not reproducible then your library balance will be way off, whatever else you do well. When troubleshooting an imbalanced pool, I recommend you repeat quantification on your individual libraries a second time, and see if you receive the same result.

It is worth asking your NGS provider to share their quantification results with you, so you can compare them to your own expectation. No two quantification measurement will ever be in precise agreement, but your NGS provider must have a very robust process in order to provide you with a reliable per-lane yield, so you can use their result as a gold-standard during troubleshooting.


If you are quantifying by qPCR, here are some valuable tips to improve robustness:


  • Perform quantification measurements in triplicate on your plate
  • Check your triplicate measurements are within ~0.5 Ct values
  • Take the Median value of your triplicates
  • Quantify all libraries which you plan to pool together on a single qPCR plate
  • Always run a no-template control to check for nonspecific amplification or contamination

If you are quantifying by qubit or bioanalyzer, I recommend that you swap to qPCR as soon as possible - and I bet you will see a better pooling balance afterwards.


Finally, have a look at #3: Was the process of mixing the libraries robust?

A common mistake when pooling is to quantify your library, perform a dilution, and then assume the diluted library will be exactly the concentration you aimed for. Unfortunately this is only true if your original concentration is close to your goal. As a guideline, any dilution greater than 1:5 is unlikely to be sufficiently robust for multiplexing. Using small volumes during dilution steps can really exacerbate this problem

The best practice for diluting highly concentrated libraries prior to pooling is to dilute them to a low value just higher than your goal, then re-quantify, then do a final small dilution to reach your goal. Use large volumes for your dilution steps, and keep your final dilution step as small as possible - and definitely less than 1:5. I often aim for a final 1:2 dilution step.


Consider this simple example:


  • Library A is at 100nM, so I dilute 1ul in 9ul of buffer to give me 10nM
  • Library B is at 300nM, so I dilute 1ul in 29ul of buffer to give me 10nM
  • Library C is at 600nM, so I dilute 1ul in 59ul of buffer to give me 10nM
  • I then mix 10ul of the diluted A, B and C. 

Frankly, my pooling balance is going to be rubbish.


Here's what I should do instead:


  • Library A is at 100nM, so I dilute 10ul in 40ul of buffer to aim for 20nM, then I re-quantify and find out it is actually at 18nM. I mix 10ul of this with 8ul of buffer to give 10nM
  • Library B is at 300nM, so I dilute 10ul in 140ul of buffer to aim for 20nM, then I re-quantify and find out it is actually at 22nM. I mix 10ul of this with 12ul of buffer to give me 10nM
  • Library C is at 600nM, so I dilute 10ul in 290ul of buffer to aim for 20nM, then I re-quantify and find out it is actually at 15nM. I mix 10ul of this with 5ul of buffer to give me 10nM
  • I then mix 10ul of the diluted A, B and C

My pooling balance will be beautiful

For the true NGS novices out there, if you don't know how I calculated the dilution steps in the example above then check this out.

If you have checked #1, #2, and #3 and everything looks perfect, then get in touch with Illumina's tech support team (techsupport@illumina.com) or with your NGS provider.

Indexing 1: A Simple NGS Pooling How-To Guide

Indexes are one of the simplest improvements in the last five years of sequencing, with the most incredible far-reaching effects. Today I will share a complementary pair of posts tackling the problems our customers experience most frequently when submitting indexed libraries for sequencing.

How do I pool my library at a defined concentration?

I get asked this a lot. Our current submission requirements are 10nM - 20nM in 15ul, but what does this mean? Is the total DNA concentration in the pool 10nM, and each individual library therefore much less? Or is it that each library within the pool is at a final concentration of 10nM?

Simply put, our submission guidelines IGNORE your indexes. Quantification and clustering cannot differentiate between indexes on a sample, so all we are interested in is the total quantity of DNA in your pool.  So, for example, if you have five libraries in a pool, the final pool DNA concentration must be at least 10nM - which means that each library within that pool is at least 2nM.

Here is the simplest at-a-glance method to dilute and pool your libraries. For more detailed hints and tips read on to my next post!
  1. Quantify and quality check all of your libraries
  2. Select a goal concentration for pooling - at or below the lowest concentration of your set of libraries.
  3. Make sure this is within our current submission guidelines.
  4. Dilute all of your libraries to that concentration, using Illumina Resuspension Buffer, EB, or 10mM Tris pH 8.5 with 0.1% Tween.
  5. Combine an equal volume of all of your libraries in your pool tube

Ta-da! You are ready to submit your pool for sequencing.


Friday, 19 September 2014

Science of the Yesteryear

Are you old enough to remember The Magic Roundabout, Hong-Kong Phooey and Mr Ben?  Did you finish your degree barely touching a computer?  When you graduated was 'genomics' a mere glint in Fred Sanger's eye?  If you answered 'yes' to any of these questions then you, like me, may feel befuddled by the dizzying speed of technological advances.

Don't despair, even when you say you went to 'Glastonbury' in 1990 and realise the app-savvy, linked-in, 'omics'-brains you are talking to weren't even born.  If you have spent years in fusty, ill-funded labs only to stumble, blinded, into the light of modern science, here are my rules for survival:

1.  Don't cry

2.  Even if you don't know what the piece of data being flashed on the screen is telling you, you are still a good person

3.  You are still making a contribution, however small

4.  Don't waste money on expensive running shoes; you'll never be able to catch the latest advances

5.  Accept your limits.  You have fewer brain cells than you had when you were 20

6.  It is inevitable that one day you will be replaced by a robot, but as the great Loudon      Wainwright III said (young folk may be more familiar with his famous son, Rufus), 'at least you've been a has-been and not just a never-was.'

7.  It's not your fault; you were born too soon.


Wednesday, 17 September 2014

qPCR quantification using our new Agilent Bravo robot

We have an exciting new instrument in the Genomics Core which will enable us to automate several of our protocols which until now have been quite labour intensive. This is the Bravo Robot by Agilent. After some previous experience with automation, I think that the results we have seen up until now are really quite promising with the aim of saving hands on time and providing consistency and accuracy within protocols.


qPCR test run - One of the first tests we ran upon installation of the Bravo was qPCR quantification. We quantify all libraries which are submitted to our sequencing service so we can aim to generate cluster densities on the flowcell which will yield large amounts of high quality data.

For this test, a single RNA-seq library was quantified using our standard method with the Illumina quantification kit by Kappa Biosystems. To test the reproducibility of the robot, we performed qPCR on this one sample 24 times as this is the maximum number of tubes which can be loaded per run.

In addition to this, each of the same 24 aliquots of this one library was set up in the same way but manually. This test was useful to determine how good the liquid handling on the Bravo is by looking at the reproducibility of the 24 replicates but additionally for comparing that between manual versus automated set up. The library had also been quantified previously so we expected the concentration to be 50nM.
Agilent Bravo Robot

The results show that there is higher variation in concentrations achieved manually although both methods slightly overestimated the expected concentration of 50nM. We saw an average concentration of 55.6nM manually (an 11% increase from 50nM) in comparison to a concentration of 52.7nM (a 5.5% increase from 50nM) on the Bravo.
Despite there being an about 1.8% difference between the average concentrations seen on the manual set-up in comparison to that of the automated, we can see that the Bravo has yielded more consistant results which is what we would expect and also good news. Since this test, we are now quantifying all SLX library submissions using the Bravo.

Although the robot will not be for general use and we will be unable to run individuals qPCR, we will be using it for all qPCR quantification for libraries submitted to us for our NGS service and for generating RNA-seq libraries. Once these protocols become robust within our lab, we will explore the option of utilising the robot in other protocols including Exome libary prep. It is additionally going to be used for automation of ChIP by the Odom group.