<!DOCTYPE art SYSTEM 'http://www.biomedcentral.com/xml/article.dtd'>
<art>
	<ui>1471-2164-15-316</ui>
	<ji>1471-2164</ji>
	<fm>
		<dochead>Database</dochead>
		<bibl>
			<title>
				<p>A customized Web portal for the genome of the ctenophore <it>Mnemiopsis leidyi</it>
				</p>
			</title>
			<aug>
				<au id="A1"><snm>Moreland</snm><mnm>Travis</mnm><fnm>R</fnm><insr iid="I1"/><email>travism@mail.nih.gov</email></au>
				<au id="A2"><snm>Nguyen</snm><fnm>Anh-Dao</fnm><insr iid="I1"/><email>nguyenan@mail.nih.gov</email></au>
				<au id="A3"><snm>Ryan</snm><mi>F</mi><fnm>Joseph</fnm><insr iid="I1"/><insr iid="I2"/><email>joseph.ryan@whitney.ufl.edu</email></au>
				<au id="A4"><snm>Schnitzler</snm><mi>E</mi><fnm>Christine</fnm><insr iid="I1"/><email>christine.schnitzler@nih.gov</email></au>
				<au id="A5"><snm>Koch</snm><mi>J</mi><fnm>Bernard</fnm><insr iid="I1"/><email>bernard.koch@nih.gov</email></au>
				<au id="A6"><snm>Siewert</snm><fnm>Katherine</fnm><insr iid="I1"/><email>ksiewert@andrew.cmu.edu</email></au>
				<au id="A7"><snm>Wolfsberg</snm><mi>G</mi><fnm>Tyra</fnm><insr iid="I1"/><email>tyra@nhgri.nih.gov</email></au>
				<au id="A8" ca="yes"><snm>Baxevanis</snm><mi>D</mi><fnm>Andreas</fnm><insr iid="I1"/><email>andy@mail.nih.gov</email></au>
			</aug>
			<insg>
				<ins id="I1"><p>Genome Technology Branch, Division of Intramural Research, National Human Genome Research Institute, National Institutes of Health, 50 South Drive, Bethesda, MD 20892, USA</p></ins>
				<ins id="I2"><p>Whitney Laboratory for Marine Bioscience, University of Florida, St. Augustine, FL 32080, USA</p></ins>
			</insg>
			<source>BMC Genomics</source>
			<section><title><p>Multicellular invertebrate genomics</p></title></section><issn>1471-2164</issn>
			<pubdate>2014</pubdate>
			<volume>15</volume>
			<issue>1</issue>
			<fpage>316</fpage>
			<url>http://www.biomedcentral.com/1471-2164/15/316</url>
			<xrefbib><pubidlist><pubid idtype="doi">10.1186/1471-2164-15-316</pubid><pubid idtype="pmpid">24773765</pubid></pubidlist></xrefbib>
		</bibl>
		<history><rec><date><day>12</day><month>11</month><year>2013</year></date></rec><acc><date><day>31</day><month>3</month><year>2014</year></date></acc><pub><date><day>28</day><month>4</month><year>2014</year></date></pub></history>
		<cpyrt><year>2014</year><collab>Moreland et al.; licensee BioMed Central Ltd.</collab><note>This is an Open Access article distributed under the terms of the Creative Commons Attribution License (<url>http://creativecommons.org/licenses/by/2.0</url>), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly credited. The Creative Commons Public Domain Dedication waiver (<url>http://creativecommons.org/publicdomain/zero/1.0/</url>) applies to the data made available in this article, unless otherwise stated.</note></cpyrt>
		<kwdg>
			<kwd>
				<it>Mnemiopsis leidyi</it>
			</kwd>
			<kwd>Genome browser</kwd>
			<kwd>Customized Web portal</kwd>
			<kwd>Gene wiki</kwd>
		</kwdg>
		<abs>
			<sec>
				<st>
					<p>Abstract</p>
				</st>
				<sec>
					<st>
						<p>Background</p>
					</st><p>
						<it>Mnemiopsis leidyi</it> is a ctenophore native to the coastal waters of the western Atlantic Ocean. A number of studies on <it>Mnemiopsis</it> have led to a better understanding of many key biological processes, and these studies have contributed to the emergence of <it>Mnemiopsis</it> as an important model for evolutionary and developmental studies. Recently, we sequenced, assembled, annotated, and performed a preliminary analysis on the 150-megabase genome of the ctenophore, <it>Mnemiopsis</it>. This sequencing effort has produced the first set of whole-genome sequencing data on any ctenophore species and is amongst the first wave of projects to sequence an animal genome <it>de novo</it> solely using next-generation sequencing technologies.</p>
				</sec>
				<sec>
					<st>
						<p>Description</p>
					</st><p>The <it>Mnemiopsis</it> Genome Project Portal (<url>http://research.nhgri.nih.gov/mnemiopsis/</url>) is intended both as a resource for obtaining genomic information on <it>Mnemiopsis</it> through an intuitive and easy-to-use interface and as a model for developing customized Web portals that enable access to genomic data. The scope of data available through this Portal goes well beyond the sequence data available through GenBank, providing key biological information not available elsewhere, such as pathway and protein domain analyses; it also features a customized genome browser for data visualization.</p>
				</sec>
				<sec>
					<st>
						<p>Conclusions</p>
					</st><p>We expect that the availability of these data will allow investigators to advance their own research projects aimed at understanding phylogenetic diversity and the evolution of proteins that play a fundamental role in metazoan development. The overall approach taken in the development of this Web site can serve as a viable model for disseminating data from whole-genome sequencing projects, framed in a way that best-serves the specific needs of the scientific community.</p>
				</sec>
			</sec>
		</abs>
	</fm>
	<bdy>
		<sec>
			<st>
				<p>Background</p>
			</st><p>Ctenophores are an important group of early-branching metazoans that are essential for understanding the evolution of multicellular animals, the relationship between genomic complexity and morphological complexity, and the molecular basis for the evolution of novel cell types such as epithelia, neurons, muscle, and stem cells. One ctenophore species that has received particular attention is <it>Mnemiopsis leidyi</it>, which is native to the coastal waters of the Atlantic Ocean. Studies in <it>Mnemiopsis</it> have advanced our understanding of a number of important biological processes such as regeneration, axial patterning, and bioluminescence <abbrgrp>
					<abbr bid="B1">1</abbr>
					<abbr bid="B2">2</abbr>
					<abbr bid="B3">3</abbr>
				</abbrgrp>. As such, <it>Mnemiopsis</it> has emerged as an important model organism for understanding the immense diversity and complexity seen in the early evolution of animals.</p><p>Despite the importance of <it>Mnemiopsis</it> as an emerging model organism, there were no high-quality genome-scale sequence data available for any ctenophore species until recently. To address this dearth of genome-scale sequence data, we recently completed the sequencing, assembly, annotation, and preliminary analysis of the 150-megabase genome of <it>Mnemiopsis leidyi</it>
				<abbrgrp>
					<abbr bid="B4">4</abbr>
				</abbrgrp>; these data will serve as an invaluable resource for the growing community of developmental, evolutionary, and marine biologists studying important questions regarding early branching metazoan biology. Initial studies utilizing these sequence data have contributed to our understanding of the evolution of gene families <abbrgrp>
					<abbr bid="B5">5</abbr>
					<abbr bid="B6">6</abbr>
					<abbr bid="B7">7</abbr>
				</abbrgrp>, signaling pathways <abbrgrp>
					<abbr bid="B8">8</abbr>
					<abbr bid="B9">9</abbr>
				</abbrgrp>, protein domains <abbrgrp>
					<abbr bid="B10">10</abbr>
				</abbrgrp>, miRNAs <abbrgrp>
					<abbr bid="B11">11</abbr>
				</abbrgrp>, and genes involved in the production and detection of light <abbrgrp>
					<abbr bid="B12">12</abbr>
				</abbrgrp>. The availability of these data also has provided a solid foundation for studies aimed at resolving the question of the phylogenetic position of the ctenophores <abbrgrp>
					<abbr bid="B4">4</abbr>
				</abbrgrp>.</p><p>In recent years, databases have been created to house whole-genome sequencing data from several emerging model organisms. Genomic data and annotation are typically made accessible via public genome portals at sequencing centers such as the US Department of Energy&#8217;s Joint Genome Institute (JGI) <abbrgrp>
					<abbr bid="B13">13</abbr>
				</abbrgrp> and the Broad Institute <abbrgrp>
					<abbr bid="B14">14</abbr>
				</abbrgrp>, while other groups have developed Web-based genomic database resources that provide additional analysis tools <abbrgrp>
					<abbr bid="B15">15</abbr>
				</abbrgrp> and browsing options <abbrgrp>
					<abbr bid="B16">16</abbr>
				</abbrgrp> to increase the utility of the data. Still others have implemented genomic resources that offer the scientific community access to genomic annotation and actively seek user contributions <abbrgrp>
					<abbr bid="B17">17</abbr>
				</abbrgrp>. Ideally, for each organism with a sequenced genome, there would be a single centralized resource where most (if not all) data retrieval and analysis could take place; this kind of resource would include, at a minimum, the ability to search, browse, and download sequence and annotation data, visualize genomic data via a Web-based browser tool, and encourage the active engagement of the scientific community in maintaining a wiki-style resource for capturing supplementary annotations of gene models and predicted proteins. Moreover, this kind of resource would (and should) be developed and maintained by researchers in the model organism community who intimately understand the needs of their scientific colleagues, presenting the data in an intuitive, user-friendly and concise manner despite its sheer volume and complexity.</p><p>In our own experience advising groups who have undertaken whole-genome sequencing projects, we have found that many of these groups do not have ready access to the kind of programming resources needed to implement some of the more &#8220;advanced&#8221; database solutions currently available. With that in mind, and to facilitate the creation of the kind of centralized genomic data resource described above, we set out to develop a generalized framework that strikes a reasonable balance between ease of implementation and documented structure, without the additional constraints posed by some publicly available database schemas.</p><p>Here, we describe the development and features of a comprehensive Web-based data portal for navigating the recently completed genome sequence of <it>Mnemiopsis leidyi</it> (<url>http://research.nhgri.nih.gov/mnemiopsis/</url>). The <it>Mnemiopsis</it> Genome Project Portal (or &#8220;MGP Portal&#8221;) is a biologist-centric resource designed with a particular emphasis on usability, intuitive navigation, and clarity. Some key features of the MGP Portal include the ability to retrieve selected nucleotide and protein sequences, the availability of whole-genome datasets for download, an integrated BLAST utility for sequence comparisons, a genome browser tool, a gene-centric wiki, and &#8220;phylogenetically informed&#8221; gene ortholog clusters mapped to human KEGG pathways. Furthermore, the scope of data accessible through this Web site goes well beyond the sequence data available at GenBank, providing other key biological information such as Gene Ontology term assignments and data from pathway and protein domain analyses. In addition, we offer a set of Perl modules that can be utilized by other scientists as a generalized framework for implementing a gene page in MediaWiki, as well as a customizable genome browser for visualizing large-scale genomic data using JBrowse.</p>
		</sec>
		<sec>
			<st>
				<p>Construction and content</p>
			</st>
			<sec>
				<st>
					<p>Displaying customized genome annotation</p>
				</st><p>A major goal during the development of the MGP Portal was to create a reproducible workflow for the creation of well-annotated and visually accessible genome repositories (Figure&#160;<figr fid="F1">1</figr>). To that end, we have adopted the JavaScript-based JBrowse <abbrgrp>
						<abbr bid="B18">18</abbr>
					</abbrgrp> as the engine for our <it>Mnemiopsis</it> Genome Browser, resulting in a clean and responsive user interface for viewing the genome assembly, gene models, and all supporting data. We have developed a number of tools in the Perl scripting language to convert sequence and annotation data into a format accepted by JBrowse (version 1.11.1), which are included as Additional files <supplr sid="S1">1</supplr>, <supplr sid="S2">2</supplr> and <supplr sid="S3">3</supplr>. In addition, we provide a Perl script that facilitates the creation of MediaWiki pages that display nucleotide and protein sequences, exonic genomic coordinates, and PFAM domains (Additional file <supplr sid="S4">4</supplr>). The actual creation of genome assemblies and annotation files is outside the scope of this manuscript but is outlined in detail elsewhere <abbrgrp>
						<abbr bid="B4">4</abbr>
					</abbrgrp>.</p>
				<fig id="F1"><title><p>Figure 1</p></title><caption><p>A visual representation of the <it>Mnemiopsis</it>&#8201;Genome Project (MGP) Portal depicting the flow of <it>Mnemiopsis</it>&#8201;sequence data into accessible internal data visualization and annotation tools</p></caption><text>
   <p><b>A visual representation of the </b><b><it>Mnemiopsis </it></b><b>Genome Project (MGP) Portal depicting the flow of </b><b><it>Mnemiopsis </it></b><b>sequence data into accessible internal data visualization and annotation tools.</b>&#8201;Colored arrows correspond to individual data types listed in the Sequence Data box (center). For example, the flow of &#8216;Assembled Transcripts&#8217; data (<it>e.g.</it>, Cufflinks- and Trinity-assembled RNA-seq transcripts) is represented by the orange arrows; these data can be viewed both as tracks in the Genome Browser and/or utilized as a BLAST database.</p>
</text><graphic file="1471-2164-15-316-1"/></fig>
				<suppl id="S1">
					<title>
						<p>Additional file 1</p>
					</title>
					<text>
						<p>
							<b>The </b><b>
								<it>scaffoldToGFF3.pl </it>
							</b><b>script reformats scaffold sequences from FASTA to GFF3.</b> The script creates a GFF3 file for each sequence in the input scaffold FASTA file and is able to handle both gapped [<it>e.g.</it>, the scaffold (SCF) track] and repeat-masked [<it>e.g.</it>, the repeat-masked (MASK) track] regions in a scaffold.</p>
					</text>
					<file name="1471-2164-15-316-S1.zip">
   <p>Click here for file</p>
</file>
				</suppl>
				<suppl id="S2">
					<title>
						<p>Additional file 2</p>
					</title>
					<text>
						<p>
							<b>The </b><b>
								<it>evmToGFF3.pl </it>
							</b><b>script parses a GFF3-formatted output file created by EvidenceModeler (EVM) by collecting data about the start and end positions of predicted genes, using this information to create well-formed GFF3 files.</b>
						</p>
					</text>
					<file name="1471-2164-15-316-S2.zip">
   <p>Click here for file</p>
</file>
				</suppl>
				<suppl id="S3">
					<title>
						<p>Additional file 3</p>
					</title>
					<text>
						<p>
							<b>The </b><b>
								<it>cufflinksToGFF3.pl </it>
							</b><b>script parses predicted transcript assemblies from the GTF-formatted file created by Cufflinks.</b>
						</p>
					</text>
					<file name="1471-2164-15-316-S3.zip">
   <p>Click here for file</p>
</file>
				</suppl>
				<suppl id="S4">
					<title>
						<p>Additional file 4</p>
					</title>
					<text>
						<p>
							<b>The </b><b>
								<it>create_wiki_page.pl </it>
							</b><b>script creates MediaWiki pages for displaying genomic data.</b> Here, we provide a sample wiki page that takes as input FASTA-formatted nucleotide and protein sequences, GFF3 files containing exonic coordinates, and an hmmscan output file containing information on Pfam-A domains. The output of this perl script is called wikipage.out.</p>
					</text>
					<file name="1471-2164-15-316-S4.zip">
   <p>Click here for file</p>
</file>
				</suppl><p>To build feature tracks, JBrowse requires properly formatted input files. While newer versions of JBrowse released subsequent to the creation of the <it>Mnemiopsis</it> Genome Browser are able to accept a slightly expanded number of inputs (including specific relational database dumps), the generic feature format version 3 (GFF3) flat file was the input type we adopted, a format that continues to be supported by JBrowse at the time of this writing (version 1.11.1). Adoption of the GFF3 format greatly facilitated the processing of multiple output file types produced by a number of data analysis programs. To begin, the <it>scaffoldToGFF3.pl</it> script (Additional file <supplr sid="S1">1</supplr>) can be used to reformat scaffold sequences from FASTA to GFF3. Two parameters (&#8211;i and &#8211;o) are required, specifying the input scaffold file and desired output directory, respectively. Two optional parameters (&#8211;l and &#8211;n) are also available. The first should be called if the gap-indication letter in the input is different from &#8216;N&#8217;&#8201;, accepting a single character (<it>e.g.</it>, X) as a substitute, while the second parameter can be called in order to specify which feature-generating program was used. The script creates a GFF3 file for each sequence in the input scaffold FASTA file and is able to handle both gapped [<it>e.g.</it>, the scaffold (SCF) track] and repeat-masked [<it>e.g.</it>, the repeat-masked (MASK) track] regions in a scaffold. Next, <it>evmToGFF3.pl</it> (Additional file <supplr sid="S2">2</supplr>) parses a GFF3-formatted output file created by EvidenceModeler (EVM) <abbrgrp>
						<abbr bid="B19">19</abbr>
					</abbrgrp> by collecting data about the start and end positions of predicted genes, using this information to create well-formed GFF3 files. The script accepts several additional parameters; the input EVM file (&#8211;i) and desired output directory (&#8211;o) must be set, while the third and optional parameter (&#8211;n) is, again, the name of the feature-generating program (<it>e.g.,</it> EVM). The third module is called <it>cufflinksToGFF3.pl</it> (Additional file <supplr sid="S3">3</supplr>) and is used to parse predicted transcript assemblies from the GTF-formatted file created by Cufflinks <abbrgrp>
						<abbr bid="B20">20</abbr>
					</abbrgrp>. The <it>cufflinksToGFF3.pl</it> script has the same behavior for transcript location as <it>evmToGFF3.pl</it> has for predicted gene location, and accepts the same three parameters (&#8211;i, &#8211;o, and &#8211;n).</p><p>To import GFF3 data into JBrowse for display as custom tracks in the main genome window, a series of three JBrowse-supplied Perl scripts (<it>prepare-refseqs.pl, biodb-to-json.pl,</it> and <it>generate-names.pl</it>) need to be executed using appropriate parameters and a system-specific configuration file, the details of which can be found in the tutorials on the JBrowse Web site (<url>http://www.jbrowse.org</url>). A number of custom CGI scripts were written to create the hyperlinks connecting features in JBrowse to the various sources of gene data.</p><p>A Perl script named <it>create_wiki_page.pl</it> has been used to create MediaWiki pages for displaying genomic data (Additional file <supplr sid="S4">4</supplr>). Here, we provide a sample wiki page that takes FASTA-formatted nucleotide and protein sequences, GFF3 files containing exonic coordinates, and an hmmscan output file containing information on Pfam-A domains as input. The <it>create_wiki_page.pl</it> script requires five parameters. The first three parameters specify a set of input files, including a nucleotide FASTA file (-n), a protein FASTA file (-a), and a Pfam-A file (-p). The remaining parameters specify a directory containing the input GFF3 files containing the exonic coordinates (-d) and an output directory (-o). In addition, we show a PHP command line that imports a wiki page into MediaWiki:<monospace>php/$WIKIHOME/maintenance/importTextFile.php -&#8201;-user=USER wikipage.out</monospace> The wikipage.out file is created by the <it>create_wiki_page.pl</it>.</p><p>Researchers interested in creating a customized genome browser or gene wiki for visualizing genome-related data are encouraged to utilize the aforementioned scripts (Additional files <supplr sid="S1">1</supplr>, <supplr sid="S2">2</supplr>, <supplr sid="S3">3</supplr> and <supplr sid="S4">4</supplr>), as their implementation satisfies the fundamental requirements of both JBrowse and MediaWiki. Furthermore, these modules may serve as a useful framework for both the development of gene wikis and more advanced genome browser tracks as new JBrowse applications are created to visualize genomic data.</p>
			</sec>
			<sec>
				<st>
					<p>User interface and genome browser implementation</p>
				</st><p>
					<it>Mnemiopsis</it> sequences are stored as individual text files, and several Perl scripts were written to retrieve these as single combined files. Multiple sequences can be downloaded via an HTML interface that was developed using a series of CGI/Perl scripts. The <it>Mnemiopsis</it> BLAST tool was implemented using ViroBLAST <abbrgrp>
						<abbr bid="B21">21</abbr>
					</abbrgrp> and runs on an Apache Web server (Additional file <supplr sid="S5">5</supplr>: Figure S1). The ViroBLAST tool was written in PHP, a server-side scripting language <abbrgrp>
						<abbr bid="B22">22</abbr>
					</abbrgrp>, and Perl <abbrgrp>
						<abbr bid="B23">23</abbr>
					</abbrgrp>, applying the stand-alone blastall program downloaded from NCBI <abbrgrp>
						<abbr bid="B24">24</abbr>
					</abbrgrp>. The <it>Mnemiopsis</it> BLAST databases were created using formatdb. PHP was used to parse the BLAST output created by ViroBLAST to generate the formatted BLAST results, including the internal links to the Genome Browser, the Gene Wiki pages and the Fetch Scaffold Tool. MediaWiki (version 1.19.11), written in PHP and Perl, was used for the Gene Wiki implementation. Perl was also used to create data for displaying on the Gene Wiki pages. The KEGG pathways pages were developed using JavaScript to display KEGG identifier, pathway, and gene symbol search lists. CGI/Perl scripts were used for KEGG pathway search functions and data downloading utilities (Additional file <supplr sid="S6">6</supplr>: Figure S2). Python (version 2.6) <abbrgrp>
						<abbr bid="B25">25</abbr>
					</abbrgrp> scripts were used to search the Pfam domains, identified using hmmscan <abbrgrp>
						<abbr bid="B26">26</abbr>
					</abbrgrp> from the HMMER suite, and CGI and JavaScript were used to display the Pfam domain search results (Additional file <supplr sid="S7">7</supplr>: Figure S3).</p>
				<suppl id="S5">
					<title>
						<p>Additional file 5: Figure S1</p>
					</title>
					<text>
						<p>The <it>Mnemiopsis</it>&#8201;BLAST tool (implemented using ViroBLAST) schematic illustrates the available user-defined input and output formats, BLAST programs, and database options. BLAST databases are provided for both <it>Mnemiopsis</it>&#8201;nucleotide (e.g., Mitochondrial genome) and protein [e.g., Protein Models (2.2)] data.</p>
					</text>
					<file name="1471-2164-15-316-S5.pdf">
   <p>Click here for file</p>
</file>
				</suppl>
				<suppl id="S6">
					<title>
						<p>Additional file 6: Figure S2</p>
					</title>
					<text>
						<p>The KEGG Pathways search function permits users to search KEGG pathways containing human genes, using a <it>Mnemiopsis</it>&#8201;homolog as the query. The relationships underling the search function are depicted as a series of associated flat files. A one-to-one relationship exists between the KEGG and PATHWAY tables and the PEP2SOURCE_ID and CLUSTERS tables. All other relationships are one&#8211;to-many or many-to-many. The CLUSTERING_ANALYSIS table is the final output representation of a KEGG Pathways query consisting of the combination of KEGG_ENTREZ_GENE, SPECIES, and CLUSTERS.</p>
					</text>
					<file name="1471-2164-15-316-S6.pdf">
   <p>Click here for file</p>
</file>
				</suppl>
				<suppl id="S7">
					<title>
						<p>Additional file 7: Figure S3</p>
					</title>
					<text>
						<p>The PFAM Domains search function parses a series of flat files illustrated here as a relational framework. The PFAM Domains schema is represented as six attributes, with connectors indicating the nature of each applicable relationship. DOMAIN_ACCESSION and DOMAIN_NAME have a one-to-one relationship. All other relationships between PFAM attributes are many-to-many.</p>
					</text>
					<file name="1471-2164-15-316-S7.pdf">
   <p>Click here for file</p>
</file>
				</suppl>
			</sec>
		</sec>
		<sec>
			<st>
				<p>Utility and discussion</p>
			</st>
			<sec>
				<st>
					<p>
						<it>Mnemiopsis</it> BLAST tool</p>
				</st><p>One feature of the <it>Mnemiopsis</it> Genome Project Portal is a customized stand-alone Web-based BLAST interface for performing nucleotide and amino acid sequence similarity searches (Additional file <supplr sid="S5">5</supplr>: Figure S1). ViroBLAST was used to implement our <it>Mnemiopsis</it> BLAST tool, producing an organized, manageable output that is easy to parse and navigate. Users may input their FASTA-formatted query sequences directly into the search box or upload sequence files from their computer. The customary set of BLAST programs is available, including BLASTN, BLASTP, BLASTX, TBLASTN, and TBLASTX. Nucleotide sequence databases include the <it>Mnemiopsis</it> genomic scaffolds (Main Scaffolds), consensus gene prediction models (Gene Models 2.2) and Unfiltered Gene Models (unincorporated predictions) described in Ryan <it>et al.</it>
					<abbrgrp>
						<abbr bid="B4">4</abbr>
					</abbrgrp>, all publicly available <it>Mnemiopsis</it> ESTs and mRNAs from GenBank (Public ESTs), the <it>Mnemiopsis</it> mitochondrial genome <abbrgrp>
						<abbr bid="B27">27</abbr>
					</abbrgrp>, Cufflinks-assembled RNA-seq transcripts, and Trinity-assembled RNA-seq transcripts <abbrgrp>
						<abbr bid="B28">28</abbr>
					</abbrgrp>. The protein sequence databases available through the MGP Portal include the translated proteins derived from the <it>Mnemiopsis</it> consensus gene prediction models (Protein Models 2.2), the unincorporated <it>Mnemiopsis</it> proteins derived from unincorporated gene prediction models (Unfiltered Protein Models), and the computationally derived <it>Mnemiopsis</it> mitochondrial proteins. Additionally, a user may create a customized user-defined BLAST database by uploading a file containing FASTA-formatted nucleotide or protein sequences of interest. BLAST output results feature customized color-coded boxes linked directly to relevant internal annotation resources, including the <it>Mnemiopsis</it> Genome Browser [B], the wiki-based <it>Mnemiopsis</it> Gene Pages [G], the Scaffold Fetch Tool [S], Unfiltered Gene Models [U], Cufflinks-assembled transcripts [C], and Trinity-assembled transcripts [T] (Figure&#160;<figr fid="F2">2</figr>).</p>
				<fig id="F2"><title><p>Figure 2</p></title><caption><p>The <it>Mnemiopsis</it>&#8201;BLAST implementation provides users an intuitive Web interface for performing sequence similarity searches</p></caption><text>
   <p><b>The </b><b><it>Mnemiopsis </it></b><b>BLAST implementation provides users an intuitive Web interface for performing sequence similarity searches.</b>&#8201;Shown are tab-delimited BLASTP results from a single human PAX3 protein, providing links to relevant sequence entries in the <it>Mnemiopsis</it>&#8201;Genome Browser (purple &#8216;B&#8217; box), individual Gene Wiki pages (green &#8216;G&#8217; box), and Unfiltered Protein Models (orange &#8216;U&#8217; box).</p>
</text><graphic file="1471-2164-15-316-2"/></fig>
			</sec>
			<sec>
				<st>
					<p>Browsing the <it>Mnemiopsis</it> genome</p>
				</st><p>One of our primary objectives in developing the MGP Portal is to provide a graphical tool for the scientific community to visualize the various types of <it>Mnemiopsis</it> genome data currently available. Using the built-in JBrowse genome browser, users can view a variety of data tracks, such as the <it>Mnemiopsis</it> genome assembly, gene prediction models, and RNA-seq data (Figure&#160;<figr fid="F3">3</figr>). Several customized JBrowse tracks are available for viewing, including consensus gene prediction models (labeled 2.2 in the browser), <it>Mnemiopsis</it> RNA-seq reads assembled into transcripts using Cufflinks (CL2) and Trinity (TRN15-30 hpf), publicly available <it>Mnemiopsis</it> ESTs from GenBank (EST), publicly available <it>Mnemiopsis</it> mRNAs from GenBank (GBNT), assembled genomic scaffolds (SCF), genomic regions that have been repeat-masked using V-Match (MASK) <abbrgrp>
						<abbr bid="B29">29</abbr>
					</abbrgrp>, experimentally verified <it>Mnemiopsis</it> RACE PCR transcripts (RACE), unincorporated gene prediction models (2.2UF), and non-redundant protein domains derived from Pfam hmmscan runs using the 2.2 and 2.2UF datasets and the six-frame translations of the <it>Mnemiopsis</it> genome (PFAM2.2).</p>
				<fig id="F3"><title><p>Figure 3</p></title><caption><p>Several customized tracks can be displayed on the <it>Mnemiopsis</it>&#8201;genome browser, implemented in JBrowse</p></caption><text>
   <p><b>Several customized tracks can be displayed on the </b><b><it>Mnemiopsis </it></b><b>genome browser, implemented in JBrowse.</b>&#8201;Here, we present a predicted protein model (ML000129a) containing G-protein-coupled receptor domains, as evidenced by the PFAM2.2 track. Transcripts assembled from RNA-seq reads support the predicted protein model as depicted in the CL2 track. Transcripts are presented as gray arrows indicating the orientation of the transcript. Exons are presented as light-colored solid bars and untranslated regions (both 5&#8217; and 3&#8217;) are darker-shaded bars. Assembled genomic scaffolds (SCF) are depicted as solid black tracks with intermittent gaps shaded bright pink. The MASK track also appears as a solid black bar highlighted with blue in the genomic regions that have been repeat-masked using VMatch. Other tracks shown include an EST and several unfiltered protein prediction models (2.2UF). Additional tracks are described in the main text.</p>
</text><graphic file="1471-2164-15-316-3"/></fig><p>The JBrowse display is organized by genomic scaffold. Scaffolds are named using a six-character convention (<it>e.g.</it>, ML<it>nnnn</it>); the ML designates the species (<it>Mnemiopsis leidyi</it>), and the individual scaffolds are numbered from 0001 to 5100 (<it>e.g.</it>, ML0001). Gene identifiers (<it>e.g.</it>, ML000129a) start with a prefix indicating the scaffold on which the gene is located (in this example, &#8216;ML0001&#8217;), followed by a non-padded integer that is unique in combination with the scaffold identifier (in this case, &#8216;29&#8217;) and ending with a lower-case letter that specifies the gene isoform (in this case, &#8216;a&#8217;). Data displayed in the browser can be searched using a variety of <it>Mnemiopsis</it> identifiers. A scaffold-based query takes the user directly to that scaffold, while a gene-based search goes to that gene&#8217;s location on the appropriate scaffold. A user may also search the genome browser using a <it>Mnemiopsis</it> GenBank mRNA identifier, an EST identifier, or a Pfam-A domain name (<it>e.g.,</it> AF293700.1, FC475136, or &#8220;Glyco hydro 20&#8221;, respectively). The PFAM2.2 track was created by running hmmscan against the Protein Models 2.2, the Unfiltered Protein Models, and the six-frame translations of the <it>Mnemiopsis</it> genome. Scaffold coordinates are displayed across the top of the browser window. Navigation options can be found beneath the coordinate bar, including the zoom tool and left-right arrows. Users can also refine the displayed region by entering the scaffold coordinates into the search box to the right of the navigation options.</p><p>Genome browser tracks are described in the Track Descriptions link above the left sidebar. The consensus <it>Mnemiopsis</it> gene prediction models (track 2.2) are presented by default. Additional track can be added by clicking on the appropriate track box on the left sidebar of the main view window. All track options are displayed in a given scaffold window even when there are no annotated features in that particular region. Features are represented as black arrows, with the direction of the arrow indicating the orientation. Exons are presented as light-colored solid bars (<it>e.g.</it>, green for 2.2 and pink for 2.2UF), while untranslated regions (both 5&#8217; and 3&#8217;) are rendered using darker-shaded colors. Assembled genomic scaffolds (SCF) are depicted as solid black tracks with intermittent bright pink gaps. The MASK track appears as a solid black bar, with blue highlighting the regions that were repeat-masked using VMatch (Figure&#160;<figr fid="F3">3</figr>). The Reference sequence track depicts the scaffold sequence, but only when the display is fully zoomed in.</p>
			</sec>
			<sec>
				<st>
					<p>
						<it>Mnemiopsis</it> genes in KEGG pathways</p>
				</st><p>Previously, we identified gene clusters that contain likely orthologs of both <it>Mnemiopsis</it> and human proteins <abbrgrp>
						<abbr bid="B4">4</abbr>
					</abbrgrp>. In order to gain some insight into the function of individual <it>Mnemiopsis</it> genes, we used this information to assign individual <it>Mnemiopsis</it> genes to human KEGG pathways. We converted Ensembl identifiers for all human genes from our phylogenetically informed clusters of orthologous genes <abbrgrp>
						<abbr bid="B4">4</abbr>
					</abbrgrp> to Entrez Gene IDs using a file from the Entrez Gene FTP site <abbrgrp>
						<abbr bid="B30">30</abbr>
					</abbrgrp>; we then used EASE <abbrgrp>
						<abbr bid="B31">31</abbr>
					</abbrgrp> to link Entrez Gene IDs to KEGG pathways. These data can be accessed by following the KEGG Pathways link found in the left sidebar of most MGP Portal pages. Human KEGG pathways containing genes with a <it>Mnemiopsis</it> ortholog are searchable by selecting a KEGG identifier (<it>e.g.,</it> hsa00604), a pathway name (<it>e.g.,</it> glycosphingolipid biosynthesis), or a gene symbol (<it>e.g.,</it> SLC33A1) from their respective lists, or by using the KEGG pathways search box (Additional file <supplr sid="S6">6</supplr>: Figure S2). The results are presented as pathway-specific ortholog cluster matrices (Figure&#160;<figr fid="F4">4</figr>).</p>
				<fig id="F4"><title><p>Figure 4</p></title><caption><p>Human KEGG pathways containing genes with a <it>Mnemiopsis</it>&#8201;homolog are presented as pathway-specific ortholog cluster matrices</p></caption><text>
   <p><b>Human KEGG pathways containing genes with a </b><b><it>Mnemiopsis </it></b><b>homolog are presented as pathway-specific ortholog cluster matrices.</b>&#8201;Each row in the glycosphingolipid biosynthesis pathway represents a cluster from our clustering analysis. The &#8216;Cluster&#8217; column indicates the most inclusive clade that encompasses all of the proteins in the cluster. The &#8216;Ratio&#8217; column represents the number of human proteins in a given cluster that are found in the pathway over the total number of human proteins in that cluster. <it>Mnemiopsis</it>&#8201;(Ml) entries are shaded in gray and hyperlinked to their respective Gene Wiki pages.</p>
</text><graphic file="1471-2164-15-316-4"/></fig><p>For each KEGG pathway, each row in the ortholog cluster table represents a cluster of orthologous genes from our clustering analysis <abbrgrp>
						<abbr bid="B4">4</abbr>
					</abbrgrp>. For a cluster to be included in the table, at least one human gene from that cluster must be present in the KEGG pathway represented. In most cases, a row consists of one or more human genes that belong to the selected KEGG pathway, along with the computationally predicted orthologs from <it>Mnemiopsis</it> and 21 other model organisms. The &#8216;Cluster&#8217; column indicates the most inclusive phylogenetic clade that encompasses all of the genes in the particular cluster (<it>e.g.,</it> &#8216;Metazoa&#8217; could indicate that the cluster contains genes from both bilaterians and cnidarians and thus cannot be characterized by a less inclusive clade, such as &#8216;Bilateria&#8217;). Each human gene is hyperlinked to its Entrez Gene entry. The &#8216;Ratio&#8217; column represents the number of human genes in the particular cluster that are in a given pathway (numerator) over the number of total human genes in that cluster (denominator). The higher the ratio, the more likely the non-human orthologs in the cluster are involved in the pathway. The numbers in the columns under each species abbreviation indicate the number of genes from that species that are in that cluster. Each number in the &#8216;Ml&#8217; (<it>Mnemiopsis leidyi</it>) column is linked to the appropriate <it>Mnemiopsis</it> gene ID(s) and corresponding Gene Wiki pages.</p>
			</sec>
			<sec>
				<st>
					<p>Pfam domains in <it>Mnemiopsis</it> proteins</p>
				</st><p>Another way to characterize the <it>Mnemiopsis</it> genes is to determine the protein domains that they encode. We used hmmscan from the HMMER suite (HMMER 3.0; March 2010) to search the Protein Models 2.2 and Unfiltered Protein Models for domains from the Pfam-A database (version 25). The gathering threshold (cut_ga) option was used to ensure conservative domain prediction. The Pfam Domains link on the home page of the MGP Portal takes the user to a query page, where researchers can specify a given Pfam-A domain by name or accession number, then search for <it>Mnemiopsis</it> genes that contain that domain (Additional file <supplr sid="S7">7</supplr>: Figure S3). The results are displayed as a list of protein models, listed by gene identifier, and the number of query domains found in those protein models. Additionally, a user may download FASTA-formatted Pfam-A domain sequences from the resulting list by clicking on the check boxes next to the sequence(s) of interest, selecting either Pfam-A domain only or the full-length domain-containing protein from the pull-down menu, and clicking &#8216;Get&#8217;.</p>
			</sec>
			<sec>
				<st>
					<p>
						<it>Mnemiopsis</it> Gene Wiki</p>
				</st><p>In an effort to engage the collective expertise of the scientific community, we have implemented a collaborative wiki (MediaWiki version 1.19.11) for the <it>Mnemiopsis</it> gene complement. The <it>Mnemiopsis</it> Gene Wiki is accessible from the left sidebar of most pages and is searchable either by selecting a <it>Mnemiopsis</it> gene identifier (<it>e.g.,</it> ML00011a) from the drop-down menu or by manually entering an identifier in the appropriate search box. Users can also access these pages by clicking on a gene in the 2.2 track of the genome browser. Each record in the Gene Wiki represents a single <it>Mnemiopsis</it> gene and provides the following annotation: nucleotide and protein sequences, coding exonic genomic coordinates, pre-computed BLAST hits from numerous organisms displaying the top hits for each protein, the top non-self BLAST hit to <it>Mnemiopsis</it>, Pfam-A domains, Gene Ontology (GO) functional annotations, human disease genes from Online Mendelian Inheritance in Man (OMIM), and a table of ortholog clusters formed by phylogenetically informed clustering methods <abbrgrp>
						<abbr bid="B4">4</abbr>
					</abbrgrp> (Figure&#160;<figr fid="F5">5</figr>). In addition, controlled editable sections have been included that permit (and encourage) the scientific community to provide further gene annotation for isoforms, <it>in situ</it> images, references, and other notes for each gene. Users interested in supplementing our gene model annotation at the <it>Mnemiopsis</it> Gene Wiki pages must first create an account and log in prior to submitting their contributions. In-house subject matter expert data curators are notified by e-mail following the creation of a new user account or an edit to an existing Gene Wiki record. Any content changes or additions to the Gene Wiki are thoroughly evaluated by these data curators and are made public subject to their approval.</p>
				<fig id="F5"><title><p>Figure 5</p></title><caption><p>Each record in the Gene Wiki represents a single <it>Mnemiopsis</it>&#8201;gene (ML000127a) and provides the following annotation: nucleotide and protein sequences, coding exonic genomic coordinates, and pre-computed BLAST hits from numerous organisms displaying the top hits for each protein</p></caption><text>
   <p><b>Each record in the Gene Wiki represents a single </b><b><it>Mnemiopsis </it></b><b>gene (ML000127a) and provides the following annotation: nucleotide and protein sequences, coding exonic genomic coordinates, and pre-computed BLAST hits from numerous organisms displaying the top hits for each protein.</b>&#8201;The inset illustrates additional annotations available through the Gene Wiki pages, including those regarding Pfam-A domains and GO terminology (functional annotation).</p>
</text><graphic file="1471-2164-15-316-5"/></fig><p>Pre-compiled BLAST hits are enumerated in tabular form. Each <it>Mnemiopsis</it> protein was compared to the UniProt and NCBI non-redundant protein databases (nr) using BLASTP. The results display the hit number, the accession numbers, <it>E</it>-values, and brief descriptions of the top four hits (lowest <it>E</it>-values). Accession numbers are linked to relevant corresponding entries at UniProt and GenBank. The <it>E</it>-values are hyperlinked to the pairwise BLAST alignments.</p><p>Each <it>Mnemiopsis</it> protein was also compared to sequence data from developmentally relevant organisms, including <it>Homo sapiens</it>, <it>Drosophila melanogaster</it>, <it>Capitella teleta</it>, <it>Amphimedon queenslandica</it>, <it>Nematostella vectensis</it>, <it>Hydra magnipapillata</it>, <it>Trichoplax adhaerens</it>, <it>Monosiga brevicollis</it>, <it>Salpingoeca rosetta</it>, <it>Capsaspora owczarzaki</it>, fungi, plants, and non-eukaryotes. The top hits for BLASTP and TBLASTN results, falling below an established <it>E</it>-value threshold (<it>E</it>-value&#8201;&#8804;&#8201;1 &#215; 10<sup>-6</sup>), are displayed along with their gene or protein identifiers, <it>E</it>-values, and description of the best hits. For the <it>Mnemiopsis</it> organismal database search, the gene identifier of the top non-self hit is displayed (and linked to its corresponding Gene Wiki page) along with the <it>E</it>-value for that alignment.</p><p>The Gene Wiki also contains a section displaying the Pfam-A domains that are encoded by the protein. The Pfam identifier, domain architectures, sequence start and end coordinates, HMM start and end coordinates, <it>E</it>-values, and domain sequence are displayed for each Pfam-A domain. All Pfam identifiers are hyperlinked to their corresponding entries on the Pfam Web site <abbrgrp>
						<abbr bid="B32">32</abbr>
					</abbrgrp>. To further assist in classification, GO terms are presented for each gene, with GO terms assigned using the Argot2 <abbrgrp>
						<abbr bid="B33">33</abbr>
					</abbrgrp> method. Functional annotations derived using Blast2GO <abbrgrp>
						<abbr bid="B34">34</abbr>
					</abbrgrp> are also presented for each gene.</p>
			</sec>
			<sec>
				<st>
					<p>Retrieving single scaffold sequences</p>
				</st><p>The genomic sequence of a single (or partial) scaffold can be retrieved using the Fetch Scaffold Tool. A user can download a FASTA-formatted full genomic scaffold sequence by following the Fetch a Scaffold link in the left sidebar of the MGP Portal home page, entering a scaffold identifier (<it>e.g.,</it> ML0001) in the search box, and selecting the Fetch sequence option (Figure&#160;<figr fid="F6">6</figr>). Partial scaffold sequences can be retrieved by adding scaffold coordinates to the above query. Alternatively, users can retrieve the reverse complement or six-frame protein translation of the scaffold by selecting the appropriate option.</p>
				<fig id="F6"><title><p>Figure 6</p></title><caption><p>The <it>Mnemiopsis</it>&#8201;Fetch Tool is used to retrieve a single or partial scaffold, its reverse complement or the six-frame protein translation</p></caption><text>
   <p><b>The </b><b><it>Mnemiopsis </it></b><b>Fetch Tool is used to retrieve a single or partial scaffold, its reverse complement or the six-frame protein translation.</b>&#8201;Displayed above is the output of a queried partial genomic scaffold for ML0001 showing the specified genomic region of interest and its six-frame protein translations (the first five translations are depicted).</p>
</text><graphic file="1471-2164-15-316-6"/></fig>
			</sec>
			<sec>
				<st>
					<p>Downloading full or partial datasets</p>
				</st><p>As noted above, one of our primary objectives in developing the MGP Portal is to simplify the dissemination of all <it>Mnemiopsis</it> sequence data to the scientific community at large. To that end, we have provided users a direct method for obtaining entire <it>Mnemiopsis</it> datasets as compressed text files. The following complete sequence datasets can be downloaded by clicking on the appropriate links located in the left sidebar: the <it>Mnemiopsis leidyi</it> genome assembly (5,100 scaffolds), the full set of <it>Mnemiopsis</it> Gene Models (16,548 genes), the <it>Mnemiopsis</it> Protein Models (16,548 proteins), the <it>Mnemiopsis</it> Unfiltered Protein Models (60,006 proteins), all publicly available <it>Mnemiopsis</it> EST sequences from NCBI (15,752 ESTs), and the full <it>Mnemiopsis</it> mitochondrial genome and protein sequences <abbrgrp>
						<abbr bid="B27">27</abbr>
					</abbrgrp> (11 proteins; Table&#160;<tblr tid="T1">1</tblr>). Optionally, users may enter a known <it>Mnemiopsis</it> sequence identifier (<it>e.g.,</it> ML0001 or ML00011a) in the search box or select from a list of identifiers to retrieve a single scaffold, gene model, protein model, or EST of interest.</p>
				<table id="T1">
					<title>
						<p>Table 1</p>
					</title>
					<caption>
						<p>
							<b>
								<it>Mnemiopsis leidyi </it>
							</b><b>complete sequence datasets available for download from the </b><b>
								<it>Mnemiopsis </it>
							</b><b>Genome Project Portal</b>
						</p>
					</caption>
					<tgroup align="left" cols="2">
						<colspec align="left" colname="c1" colnum="1" colwidth="1*"/>
						<colspec align="left" colname="c2" colnum="2" colwidth="1*"/>
						<thead valign="top">
							<row rowsep="1">
								<entry colname="c1">
									<p>
										<b>Dataset</b>
									</p>
								</entry>
								<entry colname="c2">
									<p>
										<b>Number of sequences</b>
									</p>
								</entry>
							</row>
						</thead>
						<tbody valign="top">
							<row>
								<entry colname="c1">
									<p>Genome assembly (scaffolds)</p>
								</entry>
								<entry colname="c2">
									<p>5,100</p>
								</entry>
							</row>
							<row>
								<entry colname="c1">
									<p>Gene models</p>
								</entry>
								<entry colname="c2">
									<p>16,548</p>
								</entry>
							</row>
							<row>
								<entry colname="c1">
									<p>Protein models</p>
								</entry>
								<entry colname="c2">
									<p>16,548</p>
								</entry>
							</row>
							<row>
								<entry colname="c1">
									<p>Unfiltered protein models</p>
								</entry>
								<entry colname="c2">
									<p>60,006</p>
								</entry>
							</row>
							<row>
								<entry colname="c1">
									<p>ESTs</p>
								</entry>
								<entry colname="c2">
									<p>15,752</p>
								</entry>
							</row>
							<row>
								<entry colname="c1">
									<p>Mitochondrial genome</p>
								</entry>
								<entry colname="c2">
									<p>1</p>
								</entry>
							</row>
							<row rowsep="1">
								<entry colname="c1">
									<p>Mitochondrial proteins</p>
								</entry>
								<entry colname="c2">
									<p>11</p>
								</entry>
							</row>
						</tbody>
					</tgroup>
				</table>
			</sec>
			<sec>
				<st>
					<p>Demonstrating the MGP Portal&#8217;s utility: a worked example</p>
				</st><p>The MGP Portal was developed to facilitate research that would benefit from the availability of genomic information from this emerging model organism and, to this end, it includes a number of intuitive data analysis tools. To illustrate this point, consider the case of a developmental biologist studying the human TALE class homeobox gene family (<it>e.g.,</it> PBX3; [GenBank:NP_001128250.1]) who may be interested in comparing these sequences against (or predicting novel) <it>Mnemiopsis</it> homeodomain orthologs. A straightforward approach to addressing this question would be to run a BLASTP search of the PBX3 protein sequence against the <it>Mnemiopsis</it> Protein Models (2.2) database. The <it>Mnemiopsis</it> BLAST results display a number of high-scoring candidate proteins that can be further evaluated for properties characteristic of TALE class homeodomains (<it>e.g.,</it> a TALE-type homeobox; Figure&#160;<figr fid="F1">1</figr>).</p><p>Alternatively, another biologist may be interested in searching for novel <it>Mnemiopsis</it> homeodomains, using sequence data from a closely related organism such as <it>Amphimedon</it> as the basis for their search. One approach would be to retrieve a complete set of known <it>Amphimedon</it> homeodomain proteins through an NCBI Entrez query (Search&#8201;&lt;&#8201;Protein&#8201;&gt;&#8201;for: &#8220;homeodomain AND Amphimedon[ORGN]&#8221;, which yields 31 known <it>Amphimedon</it> homeodomain proteins at the time of this writing). FASTA sequences can then be copied and pasted into the <it>Mnemiopsis</it> BLAST search window or uploaded as a file from a local computer. A unique feature of the MGP Portal BLAST implementation includes access to the Unfiltered Protein Models database, which contains the complete unincorporated (unfiltered) protein dataset derived from the <it>Mnemiopsis</it> gene prediction and annotation process. A BLASTP search against the Unfiltered Protein Models database provides additional information about possible alternate transcripts and isoforms that were screened and filtered out during the initial annotation process. These unfiltered protein models can be placed into genomic context by following the color-coded <it>Mnemiopsis</it> Genome Browser [B] links in the far right-hand column of the BLAST results. The browser provides direct access to additional annotation, including a graphical representation of the Pfam-A domain that overlaps the protein model, as well as links to the sequence of the Pfam-A domain (Figure&#160;<figr fid="F7">7</figr>).</p>
				<fig id="F7"><title><p>Figure 7</p></title><caption><p>Selecting the Genome Browser link (purple &#8216;B&#8217; box from Figure&#160;<figr fid="F1">1</figr>) from a BLASTP result entry of queried known homeodomains against the complete <it>Mnemiopsis</it>&#8201;Unfiltered Protein Models directs the user to the Genome Browser displaying the applicable <it>Mnemiopsis</it>&#8201;transcript model</p></caption><text>
   <p><b>Selecting the Genome Browser link (purple &#8216;B&#8217; box from Figure </b><figr fid="F1"> 	1</figr><b>) from a BLASTP result entry of queried known homeodomains against the complete </b><b><it>Mnemiopsis </it></b><b>Unfiltered Protein Models directs the user to the Genome Browser displaying the applicable </b><b><it>Mnemiopsis </it></b><b>transcript model.</b>&#8201;Shown above the 2.2UF track in the browser is the PFAM2.2 track displaying evidence of a homeobox in the targeted region. Clicking on the &#8216;Homeobox&#8217; link in the PFAM2.2 track opens a new browser window displaying the Pfam-A domain (ML1991_pfa) prediction results derived from a pre-compiled hmmscan run using HMMER. This Pfam-A domain record provides the genomic location of the Pfam-A domain, the genomic coordinates, transcript and protein sequences, and the hmmscan output for the homeodomain prediction.</p>
</text><graphic file="1471-2164-15-316-7"/></fig><p>Similarly, a researcher may also want to use their own custom scripts or external computational tools to further-explore the available <it>Mnemiopsis</it> data sets. In such a case, the Download Sequence links in the MGP Portal can be used to download both the Protein Models and the Unfiltered Protein Models for analysis with tools from the HMMER suite <abbrgrp>
						<abbr bid="B26">26</abbr>
					</abbrgrp> (<it>e.g.,</it> hmmsearch Homeobox.hmm ML2.2aa&#8201;&gt;&#8201;ML_novel_HDs). Predicted domains with <it>E</it>-values below an inclusion threshold (<it>e.g., E</it>-value&#8201;&lt;&#8201;0.05) could then be considered candidate homeodomains for further evaluation.</p>
			</sec>
		</sec>
		<sec>
			<st>
				<p>Conclusions</p>
			</st><p>The <it>Mnemiopsis</it> Genome Project Portal is intended as a resource for investigators from the scientific community to obtain genomic information on <it>Mnemiopsis</it> through an intuitive and easy-to-use interface; it also serves as a model for researchers undertaking the development of such a customized genome portal themselves. There are a number of comprehensive data portals available for well-established model organisms (<it>e.g.</it>, FlyBase). However, as we searched for model Web sites from which to draw inspiration for the MGP Portal, we found that many repositories for next-generation sequencing data are simply Web sites with lists of links to raw sequence data accompanied by minimal annotation, or were non-intuitive and difficult to navigate. Based on this experience, we felt that the selection and utilization of essential resources to systematically manage and disseminate the considerable amounts of data generated by these sequencing projects was imperative. Thus, the presentation and conveyance of such a genome Web portal should be intuitive, user-friendly, and concise. It is within this framework that we present the MGP Portal as such a resource for the recently completed genome sequence of <it>Mnemiopsis leidyi</it>, and we are hopeful that this resource will inspire other groups as they create Web portals of their own.</p><p>It was our intent during the development of the MGP Portal to develop a resource to maximize usability while presenting a comprehensive series of datasets not available elsewhere. Recognizing the difficulties and lessons learned from the development of such a resource, and in our continued effort to further communicate our shared experiences to the scientific community at large, we encourage other investigators to consider the proposed genome portal model and, as such, have included a series of scripts (Additional files <supplr sid="S1">1</supplr>, <supplr sid="S2">2</supplr>, <supplr sid="S3">3</supplr> and <supplr sid="S4">4</supplr>) to facilitate the conversion of output files produced by various programs. Specifically, this series of scripts can be used to format annotation data for visualization within a customized genome browser and a wiki.</p><p>As described above, the MGP Portal contains sequence-based information and several customized utilities not available elsewhere, increasing the utility of the data generated by our group in the course of our <it>Mnemiopsis</it> whole-genome sequencing project <abbrgrp>
					<abbr bid="B4">4</abbr>
				</abbrgrp>. The genome browser tool provides an intuitive interface for users to visualize the various types of data available, including data resulting from our comprehensive annotation of the <it>Mnemiopsis</it> genome. Most importantly, many features of this site make it easy for users who do not have a background in bioinformatics to straightforwardly access information presented from a comparative genomics point-of-view, without having to perform many of the analyses themselves. For instance, our phylogenetically relevant gene clusters are mapped to human KEGG pathways, providing a clear phylogenetic perspective for any particular <it>Mnemiopsis</it> gene (or pathway of genes) of interest. In addition, users may contribute to our gene annotation efforts by adding isoforms, <it>in situ</it> images, or other notes to any Gene Wiki page using a secure login. We trust that the availability of these data will allow investigators from numerous fields (such as developmental, evolutionary, and marine biology) to advance their own research projects aimed at understanding phylogenetic diversity and the evolution of proteins that play a fundamental role in metazoan development.</p>
		</sec>
		<sec>
			<st>
				<p>Availability and requirements</p>
			</st><p>The <it>Mnemiopsis</it> Genome Project Portal is freely available at <url>http://research.nhgri.nih.gov/mnemiopsis</url>, with no barriers to access. Registration is only required if users wish to contribute data to the Isoforms, <it>In situ</it> Images, References, or Notes sections of any of the Gene Wiki pages.</p>
		</sec>
		<sec>
			<st>
				<p>Competing interests</p>
			</st><p>The authors declare that they have no competing interests.</p>
		</sec>
		<sec>
			<st>
				<p>Authors&#8217; contributions</p>
			</st><p>JFR and ADB conceived the study. RTM, ADN, and TGW designed and developed the database with critical input from JFR, CES, and ADB. ADN wrote the code and implemented the interfaces and tools associated with the MGP Portal. RTM, JFR, CES, BJK, KS, and TGW performed the genome annotation and data analysis. RTM, JFR, CES, TGW, and ADB tested the Web application and tools and provided feedback. RTM and ADB wrote the manuscript, with input and suggestions from CES, JFR, and TGW. ADB directed the project. All authors read and approved the final manuscript.</p>
		</sec>
	</bdy>
	<bm>
		<ack>
			<sec>
				<st>
					<p>Acknowledgments</p>
				</st><p>This research was supported by the Intramural Research Program of the National Human Genome Research Institute, National Institutes of Health. We would like to thank Steven Bond, Mark Fredriksen, Derek Gildea, and Evan Maxwell for their thoughtful, constructive comments during the development of the Portal. We also thank Steven Bond, Derek Gildea, and Evan Maxwell for their critical reading of the manuscript.</p>
			</sec>
		</ack>
		<refgrp><bibl id="B1"><title><p>Regulation and regeneration in the ctenophore <it>Mnemiopsis leidyi</it></p></title><aug><au><snm>Henry</snm><fnm>JQ</fnm></au><au><snm>Martindale</snm><fnm>MQ</fnm></au></aug><source>Dev Biol</source><pubdate>2000</pubdate><volume>227</volume><issue>2</issue><fpage>720</fpage><lpage>733</lpage><xrefbib><pubidlist><pubid idtype="doi">10.1006/dbio.2000.9903</pubid><pubid idtype="pmpid" link="fulltext">11071786</pubid></pubidlist></xrefbib></bibl><bibl id="B2"><title><p>The Radiata and the evolutionary origins of the bilaterian body plan</p></title><aug><au><snm>Martindale</snm><fnm>MQ</fnm></au><au><snm>Finnerty</snm><fnm>JR</fnm></au><au><snm>Henry</snm><fnm>JQ</fnm></au></aug><source>Mol Phylogenet Evol</source><pubdate>2002</pubdate><volume>24</volume><issue>3</issue><fpage>358</fpage><lpage>365</lpage><xrefbib><pubidlist><pubid idtype="doi">10.1016/S1055-7903(02)00208-7</pubid><pubid idtype="pmpid" link="fulltext">12220977</pubid></pubidlist></xrefbib></bibl><bibl id="B3"><title><p>The development of bioluminescence in the ctenophore <it>Mnemiopsis leidyi</it></p></title><aug><au><snm>Freeman</snm><fnm>G</fnm></au><au><snm>Reynolds</snm><fnm>GT</fnm></au></aug><source>Dev Biol</source><pubdate>1973</pubdate><volume>31</volume><issue>1</issue><fpage>61</fpage><lpage>100</lpage><xrefbib><pubidlist><pubid idtype="doi">10.1016/0012-1606(73)90321-7</pubid><pubid idtype="pmpid" link="fulltext">4150750</pubid></pubidlist></xrefbib></bibl><bibl id="B4"><title><p>The genome of the ctenophore <it>Mnemiopsis leidyi</it> and its implications for cell type evolution</p></title><aug><au><snm>Ryan</snm><fnm>JF</fnm></au><au><snm>Pang</snm><fnm>K</fnm></au><au><snm>Schnitzler</snm><fnm>CE</fnm></au><au><snm>Nguyen</snm><fnm>AD</fnm></au><au><snm>Moreland</snm><fnm>RT</fnm></au><au><snm>Simmons</snm><fnm>DK</fnm></au><au><snm>Koch</snm><fnm>BJ</fnm></au><au><snm>Francis</snm><fnm>WR</fnm></au><au><snm>Havlak</snm><fnm>P</fnm></au><au><snm>Smith</snm><fnm>SA</fnm></au><au><snm>Putnam</snm><fnm>NH</fnm></au><au><snm>Haddock</snm><fnm>SHD</fnm></au><au><snm>Dunn</snm><fnm>CW</fnm></au><au><snm>Wolfsberg</snm><fnm>TG</fnm></au><au><snm>Mullikin</snm><fnm>JC</fnm></au><au><snm>Martindale</snm><fnm>MQ</fnm></au><au><snm>Baxevanis</snm><fnm>AD</fnm></au><au><cnm>NISC Comparative Sequencing Program</cnm></au></aug><source>Science</source><pubdate>2013</pubdate><volume>342</volume><issue>6164</issue><fpage>1242592</fpage><xrefbib><pubidlist><pubid idtype="doi">10.1126/science.1242592</pubid><pubid idtype="pmpid" link="fulltext">24337300</pubid></pubidlist></xrefbib></bibl><bibl id="B5"><title><p>The homeodomain complement of the ctenophore <it>Mnemiopsis leidyi</it> suggests that ctenophora and porifera diverged prior to the ParaHoxozoa</p></title><aug><au><snm>Ryan</snm><fnm>JF</fnm></au><au><snm>Pang</snm><fnm>K</fnm></au><au><snm>Mullikin</snm><fnm>JC</fnm></au><au><snm>Martindale</snm><fnm>MQ</fnm></au><au><snm>Baxevanis</snm><fnm>AD</fnm></au></aug><source>Evodevo</source><pubdate>2010</pubdate><volume>1</volume><issue>1</issue><fpage>9</fpage><xrefbib><pubidlist><pubid idtype="doi">10.1186/2041-9139-1-9</pubid><pubid idtype="pmcid">2959044</pubid><pubid idtype="pmpid" link="fulltext">20920347</pubid></pubidlist></xrefbib></bibl><bibl id="B6"><title><p>Nuclear receptors from the ctenophore <it>Mnemiopsis leidyi</it> lack a zinc-finger DNA-binding domain: lineage-specific loss or ancestral condition in the emergence of the nuclear receptor superfamily?</p></title><aug><au><snm>Reitzel</snm><fnm>AM</fnm></au><au><snm>Pang</snm><fnm>K</fnm></au><au><snm>Ryan</snm><fnm>JF</fnm></au><au><snm>Mullikin</snm><fnm>JC</fnm></au><au><snm>Martindale</snm><fnm>MQ</fnm></au><au><snm>Baxevanis</snm><fnm>AD</fnm></au><au><snm>Tarrant</snm><fnm>AM</fnm></au></aug><source>Evodevo</source><pubdate>2011</pubdate><volume>2</volume><issue>1</issue><fpage>3</fpage><xrefbib><pubidlist><pubid idtype="doi">10.1186/2041-9139-2-3</pubid><pubid idtype="pmcid">3038971</pubid><pubid idtype="pmpid" link="fulltext">21291545</pubid></pubidlist></xrefbib></bibl><bibl id="B7"><title><p>Evolution of sodium channels and the new view of early nervous system evolution</p></title><aug><au><snm>Liebeskind</snm><fnm>BJ</fnm></au></aug><source>Commun Integr Biol</source><pubdate>2011</pubdate><volume>4</volume><issue>6</issue><fpage>679</fpage><lpage>683</lpage><xrefbib><pubidlist><pubid idtype="pmcid">3306330</pubid><pubid idtype="pmpid">22446526</pubid></pubidlist></xrefbib></bibl><bibl id="B8"><title><p>Genomic insights into Wnt signaling in an early diverging metazoan, the ctenophore <it>Mnemiopsis leidyi</it></p></title><aug><au><snm>Pang</snm><fnm>K</fnm></au><au><snm>Ryan</snm><fnm>JF</fnm></au><au><snm>Mullikin</snm><fnm>JC</fnm></au><au><snm>Baxevanis</snm><fnm>AD</fnm></au><au><snm>Martindale</snm><fnm>MQ</fnm></au></aug><source>Evodevo</source><pubdate>2010</pubdate><volume>1</volume><issue>1</issue><fpage>10</fpage><xrefbib><pubidlist><pubid idtype="doi">10.1186/2041-9139-1-10</pubid><pubid idtype="pmcid">2959043</pubid><pubid idtype="pmpid" link="fulltext">20920349</pubid></pubidlist></xrefbib></bibl><bibl id="B9"><title><p>Evolution of the TGF-beta signaling pathway and its potential role in the ctenophore, <it>Mnemiopsis leidyi</it></p></title><aug><au><snm>Pang</snm><fnm>K</fnm></au><au><snm>Ryan</snm><fnm>JF</fnm></au><au><snm>Baxevanis</snm><fnm>AD</fnm></au><au><snm>Martindale</snm><fnm>MQ</fnm></au></aug><source>PLoS One</source><pubdate>2011</pubdate><volume>6</volume><issue>9</issue><fpage>e24152</fpage><xrefbib><pubidlist><pubid idtype="doi">10.1371/journal.pone.0024152</pubid><pubid idtype="pmcid">3169577</pubid><pubid idtype="pmpid" link="fulltext">21931657</pubid></pubidlist></xrefbib></bibl><bibl id="B10"><title><p>The diversification of the LIM superclass at the base of the metazoa increased subcellular complexity and promoted multicellular specialization</p></title><aug><au><snm>Koch</snm><fnm>BJ</fnm></au><au><snm>Ryan</snm><fnm>JF</fnm></au><au><snm>Baxevanis</snm><fnm>AD</fnm></au></aug><source>PLoS One</source><pubdate>2012</pubdate><volume>7</volume><issue>3</issue><fpage>e33261</fpage><xrefbib><pubidlist><pubid idtype="doi">10.1371/journal.pone.0033261</pubid><pubid idtype="pmcid">3305314</pubid><pubid idtype="pmpid" link="fulltext">22438907</pubid></pubidlist></xrefbib></bibl><bibl id="B11"><title><p>MicroRNAs and essential components of the microRNA processing machinery are not encoded in the genome of the ctenophore <it>Mnemiopsis leidyi</it></p></title><aug><au><snm>Maxwell</snm><fnm>EK</fnm></au><au><snm>Ryan</snm><fnm>JF</fnm></au><au><snm>Schnitzler</snm><fnm>CE</fnm></au><au><snm>Browne</snm><fnm>WE</fnm></au><au><snm>Baxevanis</snm><fnm>AD</fnm></au></aug><source>BMC Genomics</source><pubdate>2012</pubdate><volume>13</volume><issue>1</issue><fpage>714</fpage><xrefbib><pubidlist><pubid idtype="doi">10.1186/1471-2164-13-714</pubid><pubid idtype="pmcid">3563456</pubid><pubid idtype="pmpid" link="fulltext">23256903</pubid></pubidlist></xrefbib></bibl><bibl id="B12"><title><p>Genomic organization, evolution, and expression of photoprotein and opsin genes in <it>Mnemiopsis leidyi</it>: a new view of ctenophore photocytes</p></title><aug><au><snm>Schnitzler</snm><fnm>CE</fnm></au><au><snm>Pang</snm><fnm>K</fnm></au><au><snm>Powers</snm><fnm>ML</fnm></au><au><snm>Reitzel</snm><fnm>AM</fnm></au><au><snm>Ryan</snm><fnm>JF</fnm></au><au><snm>Simmons</snm><fnm>D</fnm></au><au><snm>Tada</snm><fnm>T</fnm></au><au><snm>Park</snm><fnm>M</fnm></au><au><snm>Gupta</snm><fnm>J</fnm></au><au><snm>Brooks</snm><fnm>SY</fnm></au><au><snm>Blakesley</snm><fnm>RW</fnm></au><au><snm>Yokoyama</snm><fnm>S</fnm></au><au><snm>Haddock</snm><fnm>SH</fnm></au><au><snm>Martindale</snm><fnm>MQ</fnm></au><au><snm>Baxevanis</snm><fnm>AD</fnm></au></aug><source>BMC Biol</source><pubdate>2012</pubdate><volume>10</volume><issue>1</issue><fpage>107</fpage><xrefbib><pubidlist><pubid idtype="doi">10.1186/1741-7007-10-107</pubid><pubid idtype="pmcid">3570280</pubid><pubid idtype="pmpid" link="fulltext">23259493</pubid></pubidlist></xrefbib></bibl><bibl id="B13"><title><p>The Trichoplax genome and the nature of placozoans</p></title><aug><au><snm>Srivastava</snm><fnm>M</fnm></au><au><snm>Begovic</snm><fnm>E</fnm></au><au><snm>Chapman</snm><fnm>J</fnm></au><au><snm>Putnam</snm><fnm>NH</fnm></au><au><snm>Hellsten</snm><fnm>U</fnm></au><au><snm>Kawashima</snm><fnm>T</fnm></au><au><snm>Kuo</snm><fnm>A</fnm></au><au><snm>Mitros</snm><fnm>T</fnm></au><au><snm>Salamov</snm><fnm>A</fnm></au><au><snm>Carpenter</snm><fnm>ML</fnm></au><au><snm>Signorovitch</snm><fnm>AY</fnm></au><au><snm>Moreno</snm><fnm>MA</fnm></au><au><snm>Kamm</snm><fnm>K</fnm></au><au><snm>Grimwwod</snm><fnm>J</fnm></au><au><snm>Schmutz</snm><fnm>J</fnm></au><au><snm>Shapiro</snm><fnm>H</fnm></au><au><snm>Grigoriev</snm><fnm>IV</fnm></au><au><snm>Buss</snm><fnm>LW</fnm></au><au><snm>Schierwater</snm><fnm>B</fnm></au><au><snm>Dellaporta</snm><fnm>SL</fnm></au><au><snm>Rokhsar</snm><fnm>DS</fnm></au></aug><source>Nature</source><pubdate>2008</pubdate><volume>454</volume><issue>7207</issue><fpage>955</fpage><lpage>960</lpage><xrefbib><pubidlist><pubid idtype="doi">10.1038/nature07191</pubid><pubid idtype="pmpid" link="fulltext">18719581</pubid></pubidlist></xrefbib></bibl><bibl id="B14"><title><p>Genome Index - Origins of Multicellularity; Broad Institute</p></title><note>
   <url>http://www.broadinstitute.org/annotation/genome/multicellularity_project/GenomesIndex.html</url>
</note></bibl><bibl id="B15"><title><p>StellaBase: the nematostella vectensis genomics database</p></title><aug><au><snm>Sullivan</snm><fnm>JC</fnm></au><au><snm>Ryan</snm><fnm>JF</fnm></au><au><snm>Watson</snm><fnm>JA</fnm></au><au><snm>Webb</snm><fnm>J</fnm></au><au><snm>Mullikin</snm><fnm>JC</fnm></au><au><snm>Rokhsar</snm><fnm>D</fnm></au><au><snm>Finnerty</snm><fnm>JR</fnm></au></aug><source>Nucleic Acids Res</source><pubdate>2006</pubdate><volume>34</volume><issue>Database issue</issue><fpage>D495</fpage><lpage>D499</lpage><xrefbib><pubidlist><pubid idtype="pmcid">1347383</pubid><pubid idtype="pmpid" link="fulltext">16381919</pubid></pubidlist></xrefbib></bibl><bibl id="B16"><title><p>dictyBase: a new <it>Dictyostelium discoideum</it> genome database</p></title><aug><au><snm>Kreppel</snm><fnm>L</fnm></au><au><snm>Fey</snm><fnm>P</fnm></au><au><snm>Gaudet</snm><fnm>P</fnm></au><au><snm>Just</snm><fnm>E</fnm></au><au><snm>Kibbe</snm><fnm>WA</fnm></au><au><snm>Chisholm</snm><fnm>RL</fnm></au><au><snm>Kimmel</snm><fnm>AR</fnm></au></aug><source>Nucleic Acids Res</source><pubdate>2004</pubdate><volume>32</volume><issue>Database issue</issue><fpage>D332</fpage><lpage>D333</lpage><xrefbib><pubidlist><pubid idtype="pmcid">308872</pubid><pubid idtype="pmpid" link="fulltext">14681427</pubid></pubidlist></xrefbib></bibl><bibl id="B17"><title><p>Aiptasia Wiki</p></title><note>
   <url>http://aiptasia.cs.vassar.edu/AiptasiaWiki/</url>
</note></bibl><bibl id="B18"><title><p>JBrowse: a next-generation genome browser</p></title><aug><au><snm>Skinner</snm><fnm>ME</fnm></au><au><snm>Uzilov</snm><fnm>AV</fnm></au><au><snm>Stein</snm><fnm>LD</fnm></au><au><snm>Mungall</snm><fnm>CJ</fnm></au><au><snm>Holmes</snm><fnm>IH</fnm></au></aug><source>Genome Res</source><pubdate>2009</pubdate><volume>19</volume><issue>9</issue><fpage>1630</fpage><lpage>1638</lpage><xrefbib><pubidlist><pubid idtype="doi">10.1101/gr.094607.109</pubid><pubid idtype="pmcid">2752129</pubid><pubid idtype="pmpid" link="fulltext">19570905</pubid></pubidlist></xrefbib></bibl><bibl id="B19"><title><p>Automated eukaryotic gene structure annotation using EVidenceModeler and the program to assemble spliced alignments</p></title><aug><au><snm>Haas</snm><fnm>BJ</fnm></au><au><snm>Salzberg</snm><fnm>SL</fnm></au><au><snm>Zhu</snm><fnm>W</fnm></au><au><snm>Pertea</snm><fnm>M</fnm></au><au><snm>Allen</snm><fnm>JE</fnm></au><au><snm>Orvis</snm><fnm>J</fnm></au><au><snm>White</snm><fnm>O</fnm></au><au><snm>Buell</snm><fnm>CR</fnm></au><au><snm>Wortman</snm><fnm>JR</fnm></au></aug><source>Genome Biol</source><pubdate>2008</pubdate><volume>9</volume><issue>1</issue><fpage>R7</fpage><xrefbib><pubidlist><pubid idtype="doi">10.1186/gb-2008-9-1-r7</pubid><pubid idtype="pmcid">2395244</pubid><pubid idtype="pmpid" link="fulltext">18190707</pubid></pubidlist></xrefbib></bibl><bibl id="B20"><title><p>Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation</p></title><aug><au><snm>Trapnell</snm><fnm>C</fnm></au><au><snm>Williams</snm><fnm>BA</fnm></au><au><snm>Pertea</snm><fnm>G</fnm></au><au><snm>Mortazavi</snm><fnm>A</fnm></au><au><snm>Kwan</snm><fnm>G</fnm></au><au><snm>van Buren</snm><fnm>MJ</fnm></au><au><snm>Salzberg</snm><fnm>SL</fnm></au><au><snm>Wold</snm><fnm>BJ</fnm></au><au><snm>Pachter</snm><fnm>L</fnm></au></aug><source>Nat Biotechnol</source><pubdate>2010</pubdate><volume>28</volume><issue>5</issue><fpage>511</fpage><lpage>515</lpage><xrefbib><pubidlist><pubid idtype="doi">10.1038/nbt.1621</pubid><pubid idtype="pmcid">3146043</pubid><pubid idtype="pmpid" link="fulltext">20436464</pubid></pubidlist></xrefbib></bibl><bibl id="B21"><title><p>ViroBLAST: a stand-alone BLAST web server for flexible queries of multiple databases and user's datasets</p></title><aug><au><snm>Deng</snm><fnm>W</fnm></au><au><snm>Nickle</snm><fnm>DC</fnm></au><au><snm>Learn</snm><fnm>GH</fnm></au><au><snm>Maust</snm><fnm>B</fnm></au><au><snm>Mullins</snm><fnm>JI</fnm></au></aug><source>Bioinformatics</source><pubdate>2007</pubdate><volume>23</volume><issue>17</issue><fpage>2334</fpage><lpage>2336</lpage><xrefbib><pubidlist><pubid idtype="doi">10.1093/bioinformatics/btm331</pubid><pubid idtype="pmpid" link="fulltext">17586542</pubid></pubidlist></xrefbib></bibl><bibl id="B22"><title><p>PHP</p></title><note>
   <url>http://www.php.net</url>
</note></bibl><bibl id="B23"><title><p>The Perl Programming Language</p></title><note>
   <url>http://www.perl.org</url>
</note></bibl><bibl id="B24"><title><p>NCBI BLAST FTP site</p></title><note>
   <url>ftp://ftp.ncbi.nlm.nih.gov/blast</url>
</note></bibl><bibl id="B25"><title><p>Python Programming Language</p></title><note>
   <url>http://www.python.org</url>
</note></bibl><bibl id="B26"><title><p>HMMER web server: interactive sequence similarity searching</p></title><aug><au><snm>Finn</snm><fnm>RD</fnm></au><au><snm>Clements</snm><fnm>J</fnm></au><au><snm>Eddy</snm><fnm>SR</fnm></au></aug><source>Nucleic Acids Res</source><pubdate>2011</pubdate><volume>39</volume><issue>Web Server issue</issue><fpage>W29</fpage><lpage>W37</lpage><xrefbib><pubidlist><pubid idtype="pmcid">3125773</pubid><pubid idtype="pmpid" link="fulltext">21593126</pubid></pubidlist></xrefbib></bibl><bibl id="B27"><title><p>Extreme mitochondrial evolution in the ctenophore <it>Mnemiopsis leidyi</it></p></title><aug><au><snm>Pett</snm><fnm>W</fnm></au><au><snm>Ryan</snm><fnm>JF</fnm></au><au><snm>Pang</snm><fnm>K</fnm></au><au><snm>Martindale</snm><fnm>MQ</fnm></au><au><snm>Baxevanis</snm><fnm>AD</fnm></au><au><snm>Lavrov</snm><fnm>DV</fnm></au><au><cnm>NISC Comparative Sequencing Program</cnm></au></aug><source>Mitochondrial DNA</source><pubdate>2011</pubdate><volume>22</volume><issue>4</issue><fpage>130</fpage><lpage>142</lpage><xrefbib><pubidlist><pubid idtype="doi">10.3109/19401736.2011.624611</pubid><pubid idtype="pmcid">3313829</pubid><pubid idtype="pmpid" link="fulltext">21985407</pubid></pubidlist></xrefbib></bibl><bibl id="B28"><title><p>Full-length transcriptome assembly from RNA-Seq data without a reference genome</p></title><aug><au><snm>Grabherr</snm><fnm>MG</fnm></au><au><snm>Haas</snm><fnm>BJ</fnm></au><au><snm>Yassour</snm><fnm>M</fnm></au><au><snm>Levin</snm><fnm>JZ</fnm></au><au><snm>Thompson</snm><fnm>DA</fnm></au><au><snm>Amit</snm><fnm>I</fnm></au><au><snm>Adiconis</snm><fnm>X</fnm></au><au><snm>Fan</snm><fnm>L</fnm></au><au><snm>Raychowdhury</snm><fnm>R</fnm></au><au><snm>Zeng</snm><fnm>Q</fnm></au><au><snm>Chen</snm><fnm>Z</fnm></au><au><snm>Mauceli</snm><fnm>E</fnm></au><au><snm>Hacohen</snm><fnm>N</fnm></au><au><snm>Gnirke</snm><fnm>A</fnm></au><au><snm>Rhind</snm><fnm>N</fnm></au><au><snm>di Palma</snm><fnm>F</fnm></au><au><snm>Birren</snm><fnm>BW</fnm></au><au><snm>Nusbaum</snm><fnm>C</fnm></au><au><snm>Lindblad-Toh</snm><fnm>K</fnm></au><au><snm>Friedman</snm><fnm>N</fnm></au><au><snm>Regev</snm><fnm>A</fnm></au></aug><source>Nat Biotechnol</source><pubdate>2011</pubdate><volume>29</volume><fpage>644</fpage><lpage>652</lpage><xrefbib><pubidlist><pubid idtype="doi">10.1038/nbt.1883</pubid><pubid idtype="pmcid">3571712</pubid><pubid idtype="pmpid" link="fulltext">21572440</pubid></pubidlist></xrefbib></bibl><bibl id="B29"><title><p>REPuter: the manifold applications of repeat analysis on a genomic scale</p></title><aug><au><snm>Kurtz</snm><fnm>S</fnm></au><au><snm>Choudhuri</snm><fnm>JV</fnm></au><au><snm>Ohlebusch</snm><fnm>E</fnm></au><au><snm>Schleiermacher</snm><fnm>C</fnm></au><au><snm>Stoye</snm><fnm>J</fnm></au><au><snm>Giegerich</snm><fnm>R</fnm></au></aug><source>Nucleic Acids Res</source><pubdate>2001</pubdate><volume>29</volume><issue>22</issue><fpage>4633</fpage><lpage>4642</lpage><xrefbib><pubidlist><pubid idtype="doi">10.1093/nar/29.22.4633</pubid><pubid idtype="pmcid">92531</pubid><pubid idtype="pmpid" link="fulltext">11713313</pubid></pubidlist></xrefbib></bibl><bibl id="B30"><title><p>Entrez gene: gene-centered information at NCBI</p></title><aug><au><snm>Maglott</snm><fnm>D</fnm></au><au><snm>Ostell</snm><fnm>J</fnm></au><au><snm>Pruitt</snm><fnm>KD</fnm></au><au><snm>Tatusova</snm><fnm>T</fnm></au></aug><source>Nucleic Acids Res</source><pubdate>2005</pubdate><volume>33</volume><issue>Database issue</issue><fpage>D54</fpage><lpage>D58</lpage><xrefbib><pubidlist><pubid idtype="pmcid">539985</pubid><pubid idtype="pmpid" link="fulltext">15608257</pubid></pubidlist></xrefbib></bibl><bibl id="B31"><title><p>Identifying biological themes within lists of genes with EASE</p></title><aug><au><snm>Hosack</snm><fnm>DA</fnm></au><au><snm>Dennis</snm><fnm>G</fnm><suf>Jr</suf></au><au><snm>Sherman</snm><fnm>BT</fnm></au><au><snm>Lane</snm><fnm>HC</fnm></au><au><snm>Lempicki</snm><fnm>RA</fnm></au></aug><source>Genome Biol</source><pubdate>2003</pubdate><volume>4</volume><issue>10</issue><fpage>R70</fpage><xrefbib><pubidlist><pubid idtype="doi">10.1186/gb-2003-4-10-r70</pubid><pubid idtype="pmcid">328459</pubid><pubid idtype="pmpid" link="fulltext">14519205</pubid></pubidlist></xrefbib></bibl><bibl id="B32"><title><p>The Pfam protein families database</p></title><aug><au><snm>Finn</snm><fnm>RD</fnm></au><au><snm>Mistry</snm><fnm>J</fnm></au><au><snm>Tate</snm><fnm>J</fnm></au><au><snm>Coggill</snm><fnm>P</fnm></au><au><snm>Heger</snm><fnm>A</fnm></au><au><snm>Pollington</snm><fnm>JE</fnm></au><au><snm>Gavin</snm><fnm>OL</fnm></au><au><snm>Gunasekaran</snm><fnm>P</fnm></au><au><snm>Ceric</snm><fnm>G</fnm></au><au><snm>Forslund</snm><fnm>K</fnm></au><au><snm>Holm</snm><fnm>L</fnm></au><au><snm>Sonnhammer</snm><fnm>EL</fnm></au><au><snm>Eddy</snm><fnm>SR</fnm></au><au><snm>Bateman</snm><fnm>A</fnm></au></aug><source>Nucleic Acids Res</source><pubdate>2010</pubdate><volume>38</volume><issue>Database issue</issue><fpage>D211</fpage><lpage>D222</lpage><xrefbib><pubidlist><pubid idtype="pmcid">2808889</pubid><pubid idtype="pmpid" link="fulltext">19920124</pubid></pubidlist></xrefbib></bibl><bibl id="B33"><title><p>Argot2: a large scale function prediction tool relying on semantic similarity of weighted gene ontology terms</p></title><aug><au><snm>Falda</snm><fnm>M</fnm></au><au><snm>Toppo</snm><fnm>S</fnm></au><au><snm>Pescarolo</snm><fnm>A</fnm></au><au><snm>Lavezzo</snm><fnm>E</fnm></au><au><snm>Di Camillo</snm><fnm>B</fnm></au><au><snm>Facchinetti</snm><fnm>A</fnm></au><au><snm>Cilia</snm><fnm>E</fnm></au><au><snm>Velasco</snm><fnm>R</fnm></au><au><snm>Fontana</snm><fnm>P</fnm></au></aug><source>BMC Bioinforma</source><pubdate>2012</pubdate><volume>13</volume><issue>Suppl 4</issue><fpage>S14</fpage><xrefbib><pubid idtype="doi">10.1186/1471-2105-13-S4-S14</pubid></xrefbib></bibl><bibl id="B34"><title><p>Blast2GO: a universal tool for annotation, visualization and analysis in functional genomics research</p></title><aug><au><snm>Conesca</snm><fnm>A</fnm></au><au><snm>Gotz</snm><fnm>S</fnm></au><au><snm>Garcia-Gomez</snm><fnm>JM</fnm></au><au><snm>Terol</snm><fnm>J</fnm></au><au><snm>Talon</snm><fnm>M</fnm></au><au><snm>Robles</snm><fnm>M</fnm></au></aug><source>Bioinformatics</source><pubdate>2005</pubdate><volume>21</volume><issue>18</issue><fpage>3674</fpage><lpage>3676</lpage><xrefbib><pubidlist><pubid idtype="doi">10.1093/bioinformatics/bti610</pubid><pubid idtype="pmpid" link="fulltext">16081474</pubid></pubidlist></xrefbib></bibl></refgrp>
	</bm>
</art>