Showing posts with label genomes. Show all posts
Showing posts with label genomes. Show all posts

Thursday, November 1, 2012

A genetic cartography of humans



The Phase I paper of the 1000 genomes project has been published in Nature.  Similarly to the completion of the first draft of the human genome sequence, this work constitutes a milestone in the path to understand the complex relationships between genotype and phenotype in our species. When we had the first human sequence we had, for the first time, a broad view of what were the genetic constituents of our species, no doubt that this has served to advance our understanding in many fields related to human biology and disease. What is then the significance of having 999 genomes more? I have been asked this question by some journalists in the last days. 

If one would like to describe our species purely in genetic terms, a single genome could be a good approximation, but only that, an approximation. We know that we all differ from each other genetically, and that some of these differences explain part of the observable differences (the phenotype). What is the extent and nature of the genetic differences that exists currently?, or that even existed before in the human population?, which of these differences are important in terms of phenotypic variability, including the propensity to suffer from certain diseases?, what fraction of these differences have no important effect and can vary freely?. All these questions cannot get an answer from the analyses of a single genome, and only the comparison of a large set of genomes would serve to have a better idea of what is the genome of our species.


The analogy of a map has been used several times to illustrate how the genome sequence has helped us navigate it and has enabled dramatic improvements in how we address questions related to human biology. I think the analogy is very good, since a map in itself has only a limited scientific value, since it is, basically, a description. However, similarly to how ancient maps dramatically affected the course of history, having this maps enable unanticipated scientific discoveries. This first 1000 (1092 to be extact) genomes constitutes a first cartography of human genetic variability. Providing detailed information of what mutations occur in different populations. This map is not complete, of course, but enables a good level of resolution. The authors estimate that we now have a catalogue of more than 98% of the mutations that occur at a frequency of at least 1%. Continuing with the analogy we still miss is the specific details of how the coastal areas are: like if we would see them from very far away. This missing variability may be important, since variants involved in deleterious phenotypes (disease) are expected to be at very low frequencies. Thus the effort of improving this cartography will continue and 1500 additional genomes are planned within the consortium. In parallel, many other projects and even some from particular private persons are producing more individual genome sequences. It will be important to ensure that all these information ends up in public repository, so that this information is efficiently exploited by the scientific community.  




The 1000 paper is very descriptive but already shows some important results that have an impact on how we think about the relationships of genotypes and phenotypes. They report that an individual would carry on average 200-300 variants that affect conserved residues in non-coding sequences, and even 2-4 that have been associated to disease in other studies. All individuals sequenced are healthy and thus this result tells us about the plasticity of the genome to tolerate mutations that may be deleterious in other genetic backgrounds. There is much to learn from this and the 1000 genomes will be a useful resource for studies trying to associate genetic backgrounds with disease propensity. In addition the genome sequences carry the footprints of the recent evolution of human populations, and the level of observable variability of a site can be informative of the potential functionality. Thus the possible applications of this data are many, and as I posed to a journalist. The main scientific discovery enable by this articles yet to come.

Finally, there is one important aspect that journalists do not pay much attention. Putting together this project has been a gigantic effort and has required the development of new tools and algorithms to work with this massive amount of data. Only the coordinated efforts of many groups has made this possible.This comes at a time in which such tools are desperately needed, given the growing impact of idividual genome sequencing in medicine and other fields. Similar to how an ambitious mission to bring a rover to Mars impacts scientific development beyond the particular purpose of this mission, the tools developed by the 1000 genomes project are already playing a role in hundreds other genomics project. Thus the merit of this big consortium project is not entirely the immediate scientific discoveries- at times deceiving because they are inevitably only descriptive- but their catalytic effect on a scientific field. 









 

Saturday, September 22, 2012

Can genomics save endangered species?

 
Nowadays genomics is pervading many research fields in biology, and conservation biology is not an exception anymore. The Giant panda was perhaps the first organism selected for sequencing in which the primary reason was its status as an endangered species. Since then, other species have been selected for sequencing, in an effort to contribute to their conservation. To name a few: the Californian condor, the Tiger, Tasmanian devil and the Iberian lynx, are also entering the genomic era. Our group is contributing to the efforts of sequencing and analyzing the Iberian Lynx genome, an emblematic predator of our peninsula which has the dubious honor to be the most endangered feline species on the planet. With a population below 400, a fragmented and restricted distribution area and a dangerously low level of genetic diversity, its situation is rather critical. Two years ago a consortium of Spanish research groups joined forces to sequence this species' genome. 

"Candiles" the sequenced Iberian Lynx male 
 

I have been asked many times if this effort will definitely save the species, or even whether the money would not be better invested in other efforts. How can a genome help in saving an endangered species?, are we feeding unreasonable expectations on the possible role of genomics in species conservation? Although only time will tell whether such efforts will pay off, I consider that genomics can certainly provide a new, very useful angle to species conservation. In any case, genomics should be considered just as another tool towards species conservation, rather than as the definitive solution. Species are endangered because of various causes, mostly territory loss and degradation, overexploitation, and alteration of their ecological networks. It is obvious that the main focus should be given to fight the causes that triggered population drops and create the necessary conditions for the populations to recover safely. As a powerful tool to understand a species' biology, and as a way to investigate past and current population dynamics, the availability of a genome can greatly help in understanding some of the factors that may have been decisive in population decline. Having a reference genome opens the door for a closer genetic monitoring of wild populations, not only because it enables the selection of new marker genes than can be sampled in many individuals but also because it paves the way for obtaining whole genome-level population data by re-sequencing strategies. Indeed, our project includes already re-sequencing of additional individuals from the main fragmented territories occupied by the species.

Having such kind of data is key to understand gene flow among the different populations, since it will provide a better picture of the genetic pools of the different populations. This will help to better plan crosses among captive individuals -mainly those with permanent injuries that cannot be successfully released to the wild- and future releases of their progeny. This will have a direct impact in the case of the Iberian lynx, where high levels of inbreeding and low genetic diversity exposes fragmented populations to a higher rate of diseases with a genetic basis (particularly a renal disease), and a reduced potential to overcome potential infectious diseases. A better knowledge of the genetic pool of both wild and captive populations will undoubtedly help in guiding strategies to help them recover. In addition individuals and their territories could be tracked from materials such as faeces or hairs.  Other applications may be more specific for a particular endangered species, for instance in the tasmanian devil, genomics has been used to track a transmissible cancer that causes a facial tumor disease that is transmitted by biting. 

Tasmanian devil with transmissible facial tumor


Other applications of conservation genomics that go beyond the sequencing of the endangered species itself, refer to the monitoring, using similar genomics tools, of important pathogens or symbionts of endangered species. Of course all these efforts will only be of little help if the causes that drove their decline are still around. Thus there is a growing number of promising possible applications of genomics to the conservation of endangered species, some of them already at work. I expect this field to grow fast in the coming years, as a concerned scientist I am proud that my particular corner of expertise can contribute to the noble cause of helping to keep the biodiversity of our planet.

Friday, June 1, 2012

Publicly available or not?


I have always had the naive understanding that databases such as GenBank were public, and that one was free to do research on data accessed from there, and eventually publish the results. However nothing seems to be as simple as that, since many of the genomes deposited in there have not been published yet. I have experienced myself and heard from many colleagues problematic situations regarding the use of genome data taken from public databases but yet to be published. Current guidelines are open to different interpretations, and different stakeholders (editors, reviewers, users, data producers) may have entirely different and conflicting views. With the current trend we will soon have more unpublished than published genomes in public databases, so I think it is worth re-assessing the policies. Here I share some views.

Policy guidelines regarding the use of genomic sequences prior to publication are available (see NHGRI rapid data release policy http://www.genome.gov/10506376), and set reasonable rules. For instance that data producers should deposit the data publicly and should produce a paper citable for the source of the data within a short period of time. This could precede a full genome paper in which a more througough analysis is produced. Users should not take the public data to publish an analysis focused on that genome. But this situation should not be prolonged too much. The underlying idea is to reserve the opportunity to describe the main characteristics and findings to the researchers that do the effort of sequencing, assembling, and annotating a genome, while ensuring that the data serves the advancement of science by allowing other groups to perform research on the genome data as soon as it is produced. However, there are many interpretations on what possible uses of the data should be allowed. Moreover, although indicative time-frames for the preferential exploitation of the data are given (e.g. 6 months), these are only indications. In the absence of clear-cut rules, the situation is calling for conflict. With the current flow of sequencing data, we will increasingly face the situation that data produced for public use and accessible through public databases is not associated to a paper and thus unclear whether its use should require permission. In such situations one may have different interpretations on existing rules, that of the leader of the sequencing project, that of the researcher that is accessing the data, that of the agency that financed the sequencing, and even that of the editors and reviewers of papers using available but unpublished data. Below I list some undesirable situations that highlights the contradictions of the current system. These situations are not hypothetical but rather correspond to real cases that I experienced or heard from colleagues

  • Users of public databases may unadvertedly download unpublished data, specially when they use they do this at large scales. After all they are using a public repository, and it is contradictory that public databases provide data that are not usable.
  • Most genome sequencing projects are financed using public money or from agencies that require that the data is made publicly available as soon as it is produced, but this leads to the situation above, making it difficult to sequencing project leaders to know what use is being made of their data.
  • Referees may specifically ask authors to use genomes that are in databases, or simply reject a paper because it does not use this or that “publicly available” genome in the comparative analyses. In addition referees or editors may ask for evidence of a specific permission to use unpublished data.
  • Authors willing to ask for the use of an unpublished genomes may be required to explain the exact use of the data, which expose their ideas to possible direct competitors.
  • Leaders of genome projects may feel in the right to ask for authorship in exchange of data that is available on public databases.
  • Leaders of genome projects may intentionally delay the publication of the genome paper to extend the period of preferential use. They may even decide to publish partial analysis before the genome paper.
  • Some unpublished genomes are in public databases for several years, and still different interpretations are possible of whether these data could be freely used.
  • Some genomes may never be published in the form of a genome paper, because they were sequenced with a very particular purpose.
In my opinion the current situation is too ambiguous, generates conflicts and ultimately jeopardizes the advance of science. We need clear rules, rather than guidelines, and I below propose four simple rules that would simplify the process.

  • Granting agencies and sequencing centers should specify a reasonable time-frame for preferential use (6-12 months) before the data is released. This should suffice for giving the upper hand to the research team that is doing the sequencing effort, but will also force them to focus on publish inga genome paper as soon as possible.
  • During this period, sequencing projects may announce the availability of the data for restricted use, through a specific repository that can be accessed only after a specific permission is granted. This will enable use of the data from time 0.
  • Data is released to the public repositories (at least in the form of bulk download) only after that period.
  • All data in public repositories should thus be free to be used for any purpose, regardless whether a genome paper is published.

Personally, for what the activity in my lab concerns I have taken the decision that we will use any data publicly deposited in GenBank for more than a year, for any purpose other than doing a “genome paper” (of course!). I think this is in perfect agreement with the NHGRI recommendations and will definitely save us time, and worries.

Saturday, December 10, 2011

Sequencing species.... by the thousands

 When I was giving my first steps in the field of comparative genomics, there was not much to think about when deciding which genomic datasets to use: one would just take them all. With only a few dozens of genomes, mostly of bacteria, one could have everything at hand, in the local disk, just need to update every couple of months by adding one or two more...

 These times have definitely passed, and now the flow of newly sequenced genomes is... well, overwhelming (see figure below, taken from Genomes Online). This is both a blessing and a curse for us doing comparative genomics, since we have an unprecedented amount of data which enables more resolution, but we are increasingly facing novel technical and analyitical challenges.


 Just to give a taste of this avalanche of genomes from different species (projects for sequencing genomes for a given species, such as the 1000 genomes is another story) that is coming, I here list some of the projects I am aware of that aim at sequencing thousands of genomes from a given taxonomic group.

As expected, in this kind of projects it is way more easy to come up with a bold number, than to actually define the list of species that are actually going to be sequenced. At least this is what I can tell from my involvement in the i5K initiative, in which prioritisation of species to be sequenced is not simple, since usually one wants to weigh in different criteria (phylogenetic relevance, biological, economical, and clinical importance, etc).   

 I'm sure I missed some, and, in addition, there is a growing flow of genomes that are sequenced by independent groups, including my modest own group. One common weakness of this large, and small-scale initiatives is that they sometimes come with the cost for covering the genome sequencing but do not account for the necessary bioinformatics analyses to actually make sense of the data. With the sequencing costs dropping and the potential analyses becoming more complex, the actual costs of sequencing projects will more and more be on the side of the analysis beyond the assembly and annotation phases. As a result, many bioinformatics groups are streching their resources to contribute to genomics projects without getting any specific funding.

In my opinion the planning of a sequencing project should account for all the downstream phases with their associated costs. With such an approach we may end up having a handful of genomes less, but we will definitely learn more from them. 

Sunday, December 26, 2010

Marine genomes

Craig Venter's team hits again with another large-scale sequence survey. This time they report on 197 complete prokaryotic genomes from the surface ocean's plankton. Note this means roughly a 15% increase in the total number of fully-sequenced species that was available this far. What is perhaps most notable, however, is that these species belong to a largely unexplored part of earth's microbial diversity: that living in the open ocean.



They divided the sequenced genomes into two sets: 34 that seem to be ubiquitous and abundant (present in most ocean samples), and 163 that appeared only at few locations, and compared their genome contents. Widespread planktonic species seem to have reduced genomes, with several functional classes such as gene regulation, cell motility, and membrane transport highly reduced or nearly absent. 

"cryptic escape"
In light of such differences, they propose that these ubiquitous species with reduced metabolic flexibility has adapted to survive in the open ocean by "becoming invisible", that is, reducing their size and metabolism to scape from predators and survive in a poor environment.

Altenative forces driving genome reduction
It is unclear whether reduced cell density can actually serve to scape predation from viruses or plankton-grazing organisms, and this hypothesis needs further testing. This first survey only looks at big numbers, and it is for certain that future more in-depth analyses may reveal more clues on what adaptative strategies are represented by these genomes. One interesting aspect is that forces driving genome reduction here should be different from those experienced by pathogens and endosymbionts, which constitute the best studied cases of genome reduction. In contrast to pathogens and endosymbionts, which live in rich and protected environments, marine prokaryotes thrive in one of the poorest environments with respect to nutrients. Therefore it would be very enlightening to look for differences and parallelisms in these two different adaptations that resulted in streamlined genomes.