Saturday, September 22, 2012

Can genomics save endangered species?

 
Nowadays genomics is pervading many research fields in biology, and conservation biology is not an exception anymore. The Giant panda was perhaps the first organism selected for sequencing in which the primary reason was its status as an endangered species. Since then, other species have been selected for sequencing, in an effort to contribute to their conservation. To name a few: the Californian condor, the Tiger, Tasmanian devil and the Iberian lynx, are also entering the genomic era. Our group is contributing to the efforts of sequencing and analyzing the Iberian Lynx genome, an emblematic predator of our peninsula which has the dubious honor to be the most endangered feline species on the planet. With a population below 400, a fragmented and restricted distribution area and a dangerously low level of genetic diversity, its situation is rather critical. Two years ago a consortium of Spanish research groups joined forces to sequence this species' genome. 

"Candiles" the sequenced Iberian Lynx male 
 

I have been asked many times if this effort will definitely save the species, or even whether the money would not be better invested in other efforts. How can a genome help in saving an endangered species?, are we feeding unreasonable expectations on the possible role of genomics in species conservation? Although only time will tell whether such efforts will pay off, I consider that genomics can certainly provide a new, very useful angle to species conservation. In any case, genomics should be considered just as another tool towards species conservation, rather than as the definitive solution. Species are endangered because of various causes, mostly territory loss and degradation, overexploitation, and alteration of their ecological networks. It is obvious that the main focus should be given to fight the causes that triggered population drops and create the necessary conditions for the populations to recover safely. As a powerful tool to understand a species' biology, and as a way to investigate past and current population dynamics, the availability of a genome can greatly help in understanding some of the factors that may have been decisive in population decline. Having a reference genome opens the door for a closer genetic monitoring of wild populations, not only because it enables the selection of new marker genes than can be sampled in many individuals but also because it paves the way for obtaining whole genome-level population data by re-sequencing strategies. Indeed, our project includes already re-sequencing of additional individuals from the main fragmented territories occupied by the species.

Having such kind of data is key to understand gene flow among the different populations, since it will provide a better picture of the genetic pools of the different populations. This will help to better plan crosses among captive individuals -mainly those with permanent injuries that cannot be successfully released to the wild- and future releases of their progeny. This will have a direct impact in the case of the Iberian lynx, where high levels of inbreeding and low genetic diversity exposes fragmented populations to a higher rate of diseases with a genetic basis (particularly a renal disease), and a reduced potential to overcome potential infectious diseases. A better knowledge of the genetic pool of both wild and captive populations will undoubtedly help in guiding strategies to help them recover. In addition individuals and their territories could be tracked from materials such as faeces or hairs.  Other applications may be more specific for a particular endangered species, for instance in the tasmanian devil, genomics has been used to track a transmissible cancer that causes a facial tumor disease that is transmitted by biting. 

Tasmanian devil with transmissible facial tumor


Other applications of conservation genomics that go beyond the sequencing of the endangered species itself, refer to the monitoring, using similar genomics tools, of important pathogens or symbionts of endangered species. Of course all these efforts will only be of little help if the causes that drove their decline are still around. Thus there is a growing number of promising possible applications of genomics to the conservation of endangered species, some of them already at work. I expect this field to grow fast in the coming years, as a concerned scientist I am proud that my particular corner of expertise can contribute to the noble cause of helping to keep the biodiversity of our planet.

Thursday, June 28, 2012

wrap-up of the orthology, paralogy, and function symposium at SMBE 2012

I promised some people to write a short summary of the symposium that Matthew Hahn, Marc Robinson-Rechavi, Iddo Friedberg, and I co-organized at SMBE 2012. I particularly enjoyed the symposium and the room was pretty full all the time, despite running in parallel to other interesting topics. I will just write an overall summary without going into too much details of each of the talks, and at the end I would list a number of papers that were commented on the various talks. I have to clarify that this informal wrap-up only contains my own views and has not been consensuated among the organizers. I invite any of the attendants to add comments to highlight some important aspects that I may have missed.

I’ll start by providing a summary of how all this started... which is a rather unusual way, I believe. Indeed the idea of the symposium was born in the blogosphere, in the popular Jonathan Eisen’s Tree of Life blog, where he invited Matthew Hahn to write a special guest post on the “history behind” his paper on testing the orthology conjecture. One of the conclusions from that paper was that paralogous sequences were more similar in function (and in expression patterns) than paralogs, which contradicted one of the major expectations (and assumptions) behind the theories of duplication-driven functional divergence and the strategies for inferring functions from orthologous sequences. That paper had already caused a bit of a turmoil in the orthology community (I remember this was a hot discussion during the last Quest for Orthologs meeting, at Cambridge), and several concerns were being raised about the suitability of comparisons of functional annotations from different species, and the conclusions derived within the paper. Rather rapidly, several people commented on Matt’s post and a lively discussion started (more than 40 comments in total!). The discussion was so interesting that Marc Robinson-Rechavi suggested we should bring this scientific debate in the form of a symposium in one of the upcoming conference, and so is how some of us started to work on this idea.To me it was the first time that I met the other organizers in person.

The symposium started with Eugene Koonin, who nicely introduced the topic of what conjectures could be implied by the definition of orthology, a purely evolutionary one as introduced by Walter Fitch in 1970. He then showed results from his lab that indicate that conjectures tend to hold, but that there may be exception. For instance, the conjecture that orthologs should be best reciprocal hits can be broken by an accelerated evolution in one of the true orthologs, he then showed work from other groups (Sali, Sonnhammer) on the higher conservation of structure and domain architecture in orthologs as compared to paralogs. He criticized the use of GO terms by Hahn and others and argued that one should at variety of data on function to test the conjecture. He presented results from his own group which show higher conservation of expression across species. He concluded that the functional conjecture still holds, although he observed that differences may not be spectacular.  Catherina Gushanski was next talking on changes in gene expression following segmental duplications in mammals. They have produced an impressive dataset of expression from  different tissues in various mammal species. She used that set to ask the question whether duplication was contributing more to divergence than time alone and showed that levels of expression were decreasing in younger duplicates, changes were different across different tissues. She observed no differences between one-to-one orthologs or old duplicate pairs, she also found no differences in terms of tissue specificity in orthologs vs paralogs.  Next on stage was Nicholas Furnham who presented new implementations in FUNTREE that would allow exploring functional evolution on trees. He warned that EC classification is not univocal and that can also have problems for functional comparisons. They have developed “EC-Blast” which directly measures distances between enzymatic reaction based on the molecular structures of substrates and products. Christophe Dessimoz presented results from his recent paper in which they show important biases in GO term annotations, genes from the same species and families tend to be annotated with more similar terms because of experimental biases and author biases. When correcting for this biases, the conjecture still holds. However he admitted that differences were not very big, but still significant. Romain Studer came next. He measured selection and changes in structural stability in orthologs and duplicated genes. He showed that selected sites in paralogs tend to be more clustered in the structure than in orthologs, however he observed no differences in the evolution of stability between orthologs and paralogues. He concluded that differences between paralogues may be smaller than previously thought.

After the coffee break Jianzhi Zhang told us about his work towards probing the orthology conjecture. After giving a try, he gave up of using GO terms because of the many inconsistencies, and the biases observed. He thus reverted to interrogate for conservation of protein-protein interactions using experimentally determined interactions in various yeast species. Unfortunately the many interactions to test experimentally in duplicated proteins prevented him to show a comparison of orthologs and paralogs in this talk. Nevertheless he found that all PPIs tested for orthologs were conserved, even those that seemed not to be, were caused by possible errors in previous large-scale Yeast 2 Hybrid experiments. Alex Nguyen also showed results on the budding yeast gene duplications. They focused on a more specific aspect of function: the presence of short-conserved linear motifs in protein. They found that these were more likely to disappear/diverge after the duplication event, consistent with neo- or sub-functionalization models. We moved to Drosophila with our next speaker, Lev Yamplosky who exploited expression and genomic data from the 12 Drosophila genomes. They showed larger differences in paralogs, as compared to orthologs in rates of divergence, which were also more asymmetrical. They also found that these differences varied for fast- or slow-evolving families. Finally they could also find larger differences in paralogs in terms of expression. Then it was my turn, and I mainly showed our results on comparison of expression patterns in human and mouse. Our experimental design is different from others in that we use topological dating (not sequence divergence) to establish orthologs and paralogs of a similar age, and, second, we compared always orthologs to inter-species paralogs to get rid of species-specific biases in the comparisons. Our results support a larger divergence of paralogues as compared to orthologs in tissue pattern expression. Thanks to our experimental design we could also assess that most of the differences between paralogs were gained shortly after the duplication, linking the duplication event to a big fraction of the divergence. Our last speaker was Paul Thomas who gave an overview of what can you expect and what can you not expect from GO annotations. He also showed progress on how the consortium is trying to model functional evolution through gene families, and how these models can help in the study of the relationship between orthology, paralogy and gene function.


Thus we had a diverse set of talks, most of them focusing on the comparison of different aspects of functional evolution (GO annotations, expression, functional motifs, interactions, divergence, structure) and also using varying experimental designs and species. I would say one of the main conclusion is that GO (and even EC numbers) annotation can be misleading in our ascertainment of functional evolution. My personal view is that most talks showed results consistent with the conjecture, although the level of differences between paralogs and orthologs was sometimes small. Function can be described at multiple levels, and I would expect that functional divergence after duplications may affect only one or few of these. Thus if one experimental design focuses on one of such levels it may be expected to miss divergence in the other ones. In addition those designs that average over all levels will inevitably dilute small but important aspects of functional divergence. In conclusion this is an exciting topic and with the number and variety of groups that are now interested in the topic, I am sure that we will be closer and closer to understanding the complex relationships between orthology, paralogy and functional divergence.

Some links and  papers mentioned during the symposium (I probably miss some):

Abstracts from oral presentations in SMBE, including our symposium http://imgpublic.mci-group.com/ie/PCO/OralAbstracts_Final.pdf


Another post on the orthology conjecture 

Announcement of our symposiyum 


FunTree: a resource for exploring the functional evolution of
structurally defined enzyme superfamilies.
Furnham N, Sillitoe I, Holliday GL, Cuff AL, Rahman SA, Laskowski RA,
Orengo CA, Thornton JM.
Nucleic Acids Res. 2012 Jan;40(Database issue):D776-82
http://nar.oxfordjournals.org/content/40/D1/D776.long


Brawand, D., et. al. The evolution of gene expression levels in mammalian organs. URL

 Forslund et. al. Domain conservation architecture in orthologs

Huerta-Cepas and Gabaldón Assigning duplication events to relative temporal scales in genome-wide studies.

Nehrt et. al. Testing the Ortholog Conjecture with Comparative Functional Genomic Data from Mammals http://www.ploscompbiol.org/article/info%3Adoi%2F10.1371%2Fjournal.pcbi.1002073

Nguyen et. al. Proteome-Wide Discovery of Evolutionary Conserved Sequences in Disordered Regions http://stke.sciencemag.org/cgi/content/abstract/sigtrans;5/215/rs1
 
Peterson et. al. Evolutionary constraints on structural similarity in orthologs and paralogs

Thomas et. al. On the Use of Gene Ontology Annotations to Assess Functional Similarity among Orthologs and Paralogs: A Short Report


Large-scale analysis of orthologs and paralogs under covarion-like and
constant-but-different models of amino acid evolution.
Studer RA, Robinson-Rechavi M.
Mol Biol Evol. 2010 Nov;27(11):2618-27.
http://mbe.oxfordjournals.org/content/27/11/2618.short

How confident can we be that orthologs are similar, but paralogs differ?
Studer RA, Robinson-Rechavi M.
Trends Genet. 2009 May;25(5):210-6.
http://www.sciencedirect.com/science/article/pii/S0168952509000559

Pervasive positive selection on duplicated and nonduplicated vertebrate
protein coding genes.
Studer RA, Penel S, Duret L, Robinson-Rechavi M.
Genome Res. 2008 Sep;18(9):1393-402.
http://genome.cshlp.org/content/18/9/1393.short
 

Friday, June 1, 2012

Publicly available or not?


I have always had the naive understanding that databases such as GenBank were public, and that one was free to do research on data accessed from there, and eventually publish the results. However nothing seems to be as simple as that, since many of the genomes deposited in there have not been published yet. I have experienced myself and heard from many colleagues problematic situations regarding the use of genome data taken from public databases but yet to be published. Current guidelines are open to different interpretations, and different stakeholders (editors, reviewers, users, data producers) may have entirely different and conflicting views. With the current trend we will soon have more unpublished than published genomes in public databases, so I think it is worth re-assessing the policies. Here I share some views.

Policy guidelines regarding the use of genomic sequences prior to publication are available (see NHGRI rapid data release policy http://www.genome.gov/10506376), and set reasonable rules. For instance that data producers should deposit the data publicly and should produce a paper citable for the source of the data within a short period of time. This could precede a full genome paper in which a more througough analysis is produced. Users should not take the public data to publish an analysis focused on that genome. But this situation should not be prolonged too much. The underlying idea is to reserve the opportunity to describe the main characteristics and findings to the researchers that do the effort of sequencing, assembling, and annotating a genome, while ensuring that the data serves the advancement of science by allowing other groups to perform research on the genome data as soon as it is produced. However, there are many interpretations on what possible uses of the data should be allowed. Moreover, although indicative time-frames for the preferential exploitation of the data are given (e.g. 6 months), these are only indications. In the absence of clear-cut rules, the situation is calling for conflict. With the current flow of sequencing data, we will increasingly face the situation that data produced for public use and accessible through public databases is not associated to a paper and thus unclear whether its use should require permission. In such situations one may have different interpretations on existing rules, that of the leader of the sequencing project, that of the researcher that is accessing the data, that of the agency that financed the sequencing, and even that of the editors and reviewers of papers using available but unpublished data. Below I list some undesirable situations that highlights the contradictions of the current system. These situations are not hypothetical but rather correspond to real cases that I experienced or heard from colleagues

  • Users of public databases may unadvertedly download unpublished data, specially when they use they do this at large scales. After all they are using a public repository, and it is contradictory that public databases provide data that are not usable.
  • Most genome sequencing projects are financed using public money or from agencies that require that the data is made publicly available as soon as it is produced, but this leads to the situation above, making it difficult to sequencing project leaders to know what use is being made of their data.
  • Referees may specifically ask authors to use genomes that are in databases, or simply reject a paper because it does not use this or that “publicly available” genome in the comparative analyses. In addition referees or editors may ask for evidence of a specific permission to use unpublished data.
  • Authors willing to ask for the use of an unpublished genomes may be required to explain the exact use of the data, which expose their ideas to possible direct competitors.
  • Leaders of genome projects may feel in the right to ask for authorship in exchange of data that is available on public databases.
  • Leaders of genome projects may intentionally delay the publication of the genome paper to extend the period of preferential use. They may even decide to publish partial analysis before the genome paper.
  • Some unpublished genomes are in public databases for several years, and still different interpretations are possible of whether these data could be freely used.
  • Some genomes may never be published in the form of a genome paper, because they were sequenced with a very particular purpose.
In my opinion the current situation is too ambiguous, generates conflicts and ultimately jeopardizes the advance of science. We need clear rules, rather than guidelines, and I below propose four simple rules that would simplify the process.

  • Granting agencies and sequencing centers should specify a reasonable time-frame for preferential use (6-12 months) before the data is released. This should suffice for giving the upper hand to the research team that is doing the sequencing effort, but will also force them to focus on publish inga genome paper as soon as possible.
  • During this period, sequencing projects may announce the availability of the data for restricted use, through a specific repository that can be accessed only after a specific permission is granted. This will enable use of the data from time 0.
  • Data is released to the public repositories (at least in the form of bulk download) only after that period.
  • All data in public repositories should thus be free to be used for any purpose, regardless whether a genome paper is published.

Personally, for what the activity in my lab concerns I have taken the decision that we will use any data publicly deposited in GenBank for more than a year, for any purpose other than doing a “genome paper” (of course!). I think this is in perfect agreement with the NHGRI recommendations and will definitely save us time, and worries.

Sunday, March 25, 2012

Challenges in phylogenetic tree visualization

I recently read an excellent review by Roderic Page, on the challenges in phylogenetic tree representation and visualization. It provides an overview  on existing software and tools (although he missed our ETE package, see image below for an example of ETE's visualization features). The number and diversity of existing tools is overwhelming, but probably matches the diversity of different interests and possible applications of phylogenetic trees. One may be interested in  overlaying sequence information (see below), while other would be interested in displaying information on the geographical distribution of the species. Some may need to represent uncertainty and overly different topologies, or networks to represent transfers of genetic material, the possibilities are unlimited.



 Most importantly he mentions some of the challenges of tree visualization software such as the ability to represent huge trees and to allow interactive behavior with the user. In our group we have encountered such needs and this is the reason behind implementing more visualization features in ETE. Fortunately new technologies are offering new opportunities as well, and I enjoyed imagining the possibilities that 3D visualization and touchscreen technologies will provide to researchers. Definitely is a field to follow.

 If you are interested in the topic. I recommend this video.

Saturday, March 10, 2012

Open Letter for Research in Spain

As you surely have heard, Spain is facing a serious crisis in the context of a globalized market-economy (yes, it used to be a time when economical crisis related to something more tangible, such as a serious drought or a plague, but now one can only blame abstract fluxes of financial speculations). The new government is preparing a new budget which is predicted to include the most dramatic cuts in our history. Researchers here, who have already been hit by previous cuts (see this letter), are now embracing for the worst. 


 In this context, an open letter has been put together by the Confederation of Spanish Scientific Societies, the Federation of Young Researchers and others. I recommend you to read it (some cited figures and data are very revealing), and if yo agree with it sign it, as I just did.

 Open letter for research in Spain.

Sunday, March 4, 2012

Darwin's h-index

  I guess most scientists are nowadays familiar with the term "h-index", which is a metric of citations to your published articles. More specifically the h-index correspond to the number of articles (h) that have at least h citations. Given that this index is used by many funding agencies and by peers that evaluate you for a position or competitive grant, we all hope to see it grow year by year.



  Charles Darwin lived in completely different times, he had no need to apply for grants or positions every few years and there was no system to track citations or give a "number" to the supposed "impact" of his research.  He, nevertheless, has been absorbed by the current metrics obsession and has already an h-index, computed by google scholar. 

His magic number is 63. Will this change anyway our idea of how important was Darwin's impact to Science? or it will rather help us to put the h-index into context, and highlight the difficulty of measuring true impacts?

Wednesday, February 22, 2012

Phylogenetic Tree Challenge in Encyclopedia Of Life

 The Encyclopedia of Life initiative aims at providing an open, digital resource providing comprehensive information about the diversity of life. It has recently opened a call for teams that can provide a phylogeny-aware organization of as many scientific names as possible. This text is from the call:

A prize is offered to the individual or team that can provide a very large, phylogenetically-organized set(s) of scientific names suitable for ingestion into the Encyclopedia of Life as an alternate browsing hierarchy.  

[...]


Among other factors, the total number of uniquely named nodes, node/leaf ratios and tree height may be used to compare entries so contestants should consider how they wish to trade off strict consensus versus other methods of reflecting the state of phylogenetic knowledge.
Problems to solve include 1) how to assign labels to unnamed nodes, 2) how to fill in gaps so that the set of taxa included is as comprehensive as possible, even if trees are not fully resolved or all taxa have not been analyzed, 3) how to handle competing hypotheses, 4) how to update the hierarchy at least annually.  
The winning submission must be available to EOL and others under an acceptable CC license if it is under copyright.  The tree need not be previously published in peer-reviewed form.
 
 and more information is available here.