{"id":1181,"date":"2013-05-04T12:14:02","date_gmt":"2013-05-04T04:14:02","guid":{"rendered":"http:\/\/www.hzaumycology.com\/chenlianfu_blog\/?p=1181"},"modified":"2013-05-07T02:19:38","modified_gmt":"2013-05-06T18:19:38","slug":"pasa%e7%9a%84%e4%bd%bf%e7%94%a8","status":"publish","type":"post","link":"http:\/\/www.chenlianfu.com\/?p=1181","title":{"rendered":"PASA\u7684\u4f7f\u7528"},"content":{"rendered":"<p>\u524d\u6587\u8bb2\u8ff0\u4e86\uff1a<a href=\"http:\/\/www.hzaumycology.com\/chenlianfu_blog\/?p=1133\" target=\"_blank\">PASA\u7684\u5b89\u88c5\uff0c\u914d\u7f6e\u4e0e\u4e3b\u7a0b\u5e8f\u4f7f\u7528\u53c2\u6570<\/a>\u3002\u672c\u6587\u5219\u91cd\u70b9\u8bb2\u8ff0PASA\u7684\u51e0\u79cd\u5e73\u884c\u7684\u7528\u6cd5\uff1a<\/p>\n<p>\u6700\u7b80\u5355\u5730\u7406\u89e3PASA\u7684\u7528\u6cd5\uff1aPASA\u662f\u901a\u8fc7\u8c03\u7528 blat \u6216 gsnap \u6765\u5c06 transcripts sequences \u6bd4\u5bf9\u5230 genome \u4e0a\uff0c\u4ece\u800c\u8fdb\u884c\u57fa\u56e0\u7684\u7ed3\u6784\u6ce8\u91ca(gff\u6587\u4ef6\u5f62\u5f0f)\u3002<\/p>\n<h1>1. \u666e\u901a\u7528\u6cd5<\/h1>\n<p>\u4f7f\u7528transcripts sequences\u6765\u8fdb\u884c\u57fa\u56e0\u9884\u6d4b\uff0c\u9700\u89812\u4e2a\uff08\u62163\u4e2a\uff09\u8f93\u5165\u7684\u6587\u4ef6\uff1a<\/p>\n<pre>The genome sequence in a multiFasta file (ie. genome.fasta)\r\n\r\nThe transcript sequences in a multiFasta file (ie. transcripts.fasta)\r\n\r\nOptional: a file containing the list of accessions corresponding to full-length cDNAs (ie. FL_accs.txt)\u8be5\u6587\u4ef6\u662f\u7b2c\u4e8c\u4e2a\u6587\u4ef6\u4e2d\u5c5e\u4e8e\u5168\u957fCDNAs\u7684\u5e8f\u5217\u540d\u7684\u96c6\u5408\uff0ctxt\u6587\u6863\u3002<\/pre>\n<p>\u5176\u4f7f\u7528\u6b65\u9aa4\u4e3a\uff1a<\/p>\n<h2>1.1 \u4f7f\u7528seqclean\u6765\u5bf9transcript sequences\u8fdb\u884cclean<\/h2>\n<p>\u901a\u8fc7\u6c61\u67d3\u6570\u636e\u5e93vectors.fasta\u6765\u5bf9transcript sequences\u53ef\u80fd\u88ab\u6c61\u67d3\u7684\u6570\u636e\u8fdb\u884c\u6e05\u9664\u3002<\/p>\n<pre>$ seqclean transcripts.fasta -v vectors.fasta<\/pre>\n<h2>1.2 \u8fd0\u884cPASA<\/h2>\n<pre>$ cp $PASA_HOME\/pasa_conf\/pasa.alignAssembly.Template.txt alignAssembly.config\r\n$ perl -p -i -e 's\/MYSQLDB=.*\/MYSQLDB=sample_mydb_pasa\/' alignAssembly.config\r\n$ $PASA_HOME\/scripts\/Launch_PASA_pipeline.pl -c alignAssembly.config -C -R -g genome_sample.fasta -t all_transcripts.fasta.clean -T -u all_transcripts.fasta -f FL_accs.txt --ALIGNERS blat,gmap --CPU 2<\/pre>\n<h2>1.3 PASA\u7ed3\u679c<\/h2>\n<p><span style=\"color: #888888;\"><em>sample_mydb_pasa.assemblies.fasta<\/em><\/span> : the PASA assemblies in FASTA format.<\/p>\n<p><span style=\"color: #888888;\"><em>sample_mydb_pasa.pasa_assemblies.gff3,.gtf,.bed<\/em><\/span> : the PASA assembly structures.<\/p>\n<p><span style=\"color: #888888;\"><em>sample_mydb_pasa.pasa_alignment_assembly_building.ascii_illustrations.out<\/em><\/span> : descriptions of alignment assemblies and how they were constructed from the underlying transcript alignments.<\/p>\n<p><span style=\"color: #888888;\"><em>sample_mydb_pasa.pasa_assemblies_described.txt<\/em><\/span> : tab-delimited format describing the contents of the PASA assemblies, including the identity of those transcripts that were assembled into the corresponding structure.<\/p>\n<h1>2. PASA\u7ed3\u5408RNA-seq\u6570\u636e\u6765\u8fdb\u884c\u57fa\u56e0\u9884\u6d4b<\/h1>\n<p>\u7531RNA-seq\u6570\u636e\u53ef\u4ee5\u5f97\u5230transcript sequences\uff0c\u4ee5RNA-seq\u6570\u636e\u4f7f\u7528Trinity\u7ec4\u88c5\u8fc7\u540e\u7684\u8f6c\u5f55\u7ec4\u7ed3\u679c\u53ef\u4ee5\u6309\u7b2c\u4e00\u79cd\u65b9\u6cd5\u8fdb\u884cPASA\u5206\u6790\u3002\u5728\u6b64\uff0c\u66f4\u597d\u7684\u65b9\u6cd5\u662f\uff1a \u5229\u7528PASA\uff0c\u4f7f\u7528De novo\u8f6c\u5f55\u7ec4\u548cgenome\u6765\u5efa\u7acb\u4e00\u4e2a\u7efc\u5408\u8f6c\u5f55\u7ec4\u6570\u636e\u5e93\u3002<\/p>\n<p>\u9700\u89812\u4e2a(\u62163\u4e2a)\u7684\u8f93\u5165\u6587\u4ef6\uff1a<\/p>\n<pre>1. <a href=\"http:\/\/trinityrnaseq.sourceforge.net\/index.html\" target=\"_blank\">Trinity de novo<\/a> RNA-Seq assemblies (ex. Trinity.fasta)\r\n2. <a href=\"http:\/\/trinityrnaseq.sourceforge.net\/genome_guided_trinity.html\" target=\"_blank\">Trinity genome-guided<\/a> RNA-Seq assemblies (ex. Trinity.GG.fasta)\r\n3. (optionally)<a href=\"http:\/\/cufflinks.cbcb.umd.edu\/\" target=\"_blank\">cufflinks<\/a> transcript structures (ex. cufflinks.gtf).<\/pre>\n<p>\u4f7f\u7528\u6b65\u9aa4\uff1a<br \/>\n2.1 Concatenate the Trinity.fasta and Trinity.GG.fasta files into a single transcripts.fasta file.<\/p>\n<p>Trinity.GG.fasta\u6587\u4ef6\u53ef\u4ee5\u662f <a href=\"http:\/\/trinityrnaseq.sourceforge.net\/genome_guided_trinity.html\" target=\"_blank\">Trinity genome-guided<\/a> \u8fc7\u7a0b\u4e2d\u4ea7\u751f\u7684\u6587\u4ef6Trinity_GG.fasta, \u5982\u679c\u4f7f\u7528\u6700\u540ePASA\u5f97\u5230\u7684assembly\u7ed3\u679c\u6587\u4ef6\uff0c\u5219\u9700\u8981\u4fee\u6539fasta\u6587\u4ef6\u7684\u5934\u3002<\/p>\n<pre>cat Trinity.fasta Trinity.GG.fasta &gt; transcripts.fasta<\/pre>\n<p>2.2 Create a file containing the list of transcript accessions that correspond to the Trinity de novo assembly (full de novo, not genome-guided).<\/p>\n<pre>$PASA_HOME\/misc_utilities\/accession_extractor.pl &lt; Trinity.fasta &gt; tdn.accs<\/pre>\n<p>2.3 Run PASA using RNA-Seq related options as described in the section above, but include the parameter setting <span style=\"color: #888888;\"><em>&#8211;TDN tdn.accs<\/em><\/span>. To (optionally) include cufflinks-generated transcript structures, further include the parameter setting <span style=\"color: #888888;\"><em>&#8211;cufflinks_gtf cufflinks.gtf<\/em><\/span>. Note, cufflinks may not be appropriate for gene-dense targets, such as in fungi; cufflinks excels when applied to vertebrate genomes, so best to include when applying to mouse or human.<\/p>\n<pre>$ $PASA_HOME\/scripts\/Launch_PASA_pipeline.pl -c alignAssembly.config -t transcripts.fasta -C -R -g genome_sample.fasta --ALIGNERS blat,gmap --CPU 2 --TDN tdn.accs --cufflinks_gtf cufflinks.gtf<\/pre>\n<p>2.4 After completing the PASA alignment assembly, generate the comprehensive transcriptome database via:<\/p>\n<pre>$PASA_HOME\/PASA\/scripts\/build_comprehensive_transcriptome.dbi -c alignAssembly.config -t transcripts.fasta --min_per_ID 95 --min_per_aligned 30<\/pre>\n<p>2.5 \u7ed3\u679c\u6587\u4ef6<br \/>\n<span style=\"color: #888888;\"><em>compreh_init_build\/compreh_init_build.fasta<\/em><\/span> : the transcript sequences<br \/>\n<span style=\"color: #888888;\"><em>compreh_init_build\/compreh_init_build.geneToTrans_mapping<\/em><\/span> : the gene\/transcript mapping file (for use with RSEM, Trinotate, other tools)<\/p>\n<p><span style=\"color: #888888;\"><em>compreh_init_build\/compreh_init_build.bed<\/em><\/span> : transcript structures in bed format<br \/>\n<span style=\"color: #888888;\"><em>compreh_init_build\/compreh_init_build.gff3<\/em><\/span> : transcript structures in gff3 format<\/p>\n<p><span style=\"color: #888888;\"><em>compreh_init_build\/compreh_init_build.details<\/em><\/span> : classifications of transcripts according to genome mapping status.<br \/>\n2.6 \u4f7f\u7528\u611f\u53d7\uff1a<\/p>\n<p>\u5efa\u7acb\u4e00\u4e2a\u7efc\u5408\u7684transcritome\u6570\u636e\u5e93\uff0c\u8d77\u59cb\u662f\u6ca1\u5c06de novo\u7684transcripts\u7528\u4e8ePASA\u7684\u8fd0\u884c\uff0c\u7528\u4e8ePASA\u8fd0\u884c\u662f\u7684Trinity_GG.fasta\u3002\u6700\u540e\u7684\u7ed3\u679cfasta\u6587\u4ef6\u662f\u5355\u72ec\u4f7f\u7528Trinity_GG.fasta\u8fd0\u884cPASA\u5f97\u5230\u7684gff\u6587\u4ef6\uff0cfasta\u6587\u4ef6\u662f\u5355\u72ec\u8fd0\u884c\u5f97\u5230\u7684assembly.fasta\u6587\u4ef6\u548cTrinity.fasta\u76f8\u52a0\u7684\u7ed3\u679c\u3002\u4e2a\u4eba\u89c9\u5f97\u8fd9\u6837\u505a\u6ca1\u4ec0\u4e48\u610f\u4e49\uff01\u867d\u7136\uff0c\u6587\u6863\u5199\u9053\u8fd9\u6837\u505a\u7684\u610f\u4e49\u662f\u6709\u5229\u4e8e\u4e0b\u6e38\u57fa\u56e0\u5dee\u5f02\u8868\u8fbe\u7684\u5206\u6790\u3002\u8fd8\u4e0d\u5982\u5c31\u4f7f\u7528Trinity_GG.fasta\u5355\u72ec\u53bb\u8fd0\u884cPASA\u3002<\/p>\n<h1>3. Annotation Comparisons and Annotation Updates<\/h1>\n<p>Incorporating PASA Assemblies into Existing Gene Predictions, Changing Exons, Adding UTRs and Alternatively Spliced Models<\/p>\n<p>The PASA software can update any preexisting set of protein-coding gene annotations to incorporate the PASA alignment evidence, correcting exon boundaries, adding UTRs, and models for alternative splicing based on the PASA alignment assemblies generated above.<\/p>\n<p>\u6b65\u9aa4\u5982\u4e0b\uff1a<\/p>\n<h2>3.1 \u5c06protein-coding gene annotations\u4e0a\u8f7d\u5230\u6570\u636e\u5e93<\/h2>\n<p>\u9a8c\u8bc1gff3\u6587\u4ef6\u7684\u517c\u5bb9\u6027\uff0c\u5982\u679c\u4f1a\u62a5\u544a\u9519\u8bef\u548c\u8b66\u793a\uff0c\u5219\u9700\u8981\u4fee\u6539gff3\u6587\u4ef6.<\/p>\n<pre>$ $PASA_HOME\/misc_utilities\/pasa_gff3_validator.pl orig_annotations_sample.gff3<\/pre>\n<p>\u5c06\u57fa\u56e0\u6ce8\u91ca\u6587\u4ef6orig_annotations_sample.gff3\u4e0a\u8f7d\u5230\u6570\u636e\u5e93\u3002note that you should only feed protein-coding genes to PASA using the loader\u3002<\/p>\n<pre>$ $PASA_HOME\/scripts\/Load_Current_Gene_Annotations.dbi -c alignAssembly.config -g genome_sample.fasta -P orig_annotations_sample.gff3<\/pre>\n<p>\u7ee7\u7eed\u68c0\u67e5\u4e00\u904d\u517c\u5bb9\u6027\uff0c\u4e0a\u8f7d\u4e4b\u540e\uff0c\u5219\u4e0d\u4f1a\u6709\u4efb\u4f55\u8fd4\u56de\u5230\u6807\u51c6\u8f93\u51fa\u3002<\/p>\n<h2>3.2 Performing an annotation comparison and generating an updated gene set<\/h2>\n<p>\u8fd0\u884cPASA\u4e3b\u7a0b\u5e8f\uff0c\u4f7f\u7528\u6ce8\u91ca\u7684config\u6587\u4ef6(\u5c06\u5176\u4e2d\u7684MYSQLDB\u7684\u503c\u66f4\u6539\u7684\u548c\u6bd4\u5bf9\u7684config\u6587\u4ef6\u4e00\u81f4)<\/p>\n<pre>$PASA_HOME\/scripts\/Launch_PASA_pipeline.pl -c annotCompare.config -A -g genome_sample.fasta -t all_transcripts.fasta.clean<\/pre>\n<h2>3.3 \u8fd0\u884c\u7ed3\u679c<\/h2>\n<p>Once the annotation comparison is complete, PASA will output a new GFF3 file that contains the PASA-updated version of the genome annotation, including those gene models successfully updated by PASA, and those that remained untouched. This file will be named <em>${mysql_db}.gene_structures_post_PASA_updates.$pid.gff3<\/em>, where $pid is the process ID for this annotation comparison computation.<\/p>\n<p>\u5f53\u7136\uff0c\u4e5f\u53ef\u4ee5\u518d\u6b21\u770b\u770bPASA\u7684Web Portal.\u62a5\u544a\u5df2\u7ecf\u53d1\u751f\u4e86\u6539\u53d8\u3002<\/p>\n<h1>4. Identification and Classification of All Alternative Splicing Variations<\/h1>\n<p>\u5728PASA\u8fdb\u884ctranscripts\u6bd4\u5bf9\u5230genome\u65f6\uff0c\u53ef\u4ee5\u8fdb\u884c\u53ef\u53d8\u526a\u5207\u9884\u6d4b\u3002\u52a0\u5165\u9009\u9879&#8211;ALT_SPLICE\u3002<\/p>\n<h1>5. Extraction of ORFs from PASA assemblies (reference ORFs for training gene predictors)<\/h1>\n<p>PASA\u53ef\u4ee5\u4ece\u7528\u4e8e\u81ea\u52a8\u63d0\u53d6\u7f16\u7801\u86cb\u767d\u57fa\u56e0\u7684\u57fa\u56e0\u7ed3\u6784\uff0c\u7528\u4e8e\u8bad\u7ec3\u57fa\u56e0\u96c6\u5408\uff0c\u6765\u7528\u4e8e\u5176\u5b83\u57fa\u56e0\u9884\u6d4b\u8f6f\u4ef6\u3002To do this, the longest ORF (min 100 amino acids, configurable) is found within each PASA alignment assembly, allowing for 5&#8242; and 3&#8242; partials. The top 500 longest ORFs (configurable) are extracted and, using these in addition to the full set of genome sequences, log-likelihood scores for a hexamer-based Markov model are computed. All ORFs are scored using this Markov model, and those ORFs that are most likely coding are reported as best. The scoring method for ORF identification is similar to that used in the GeneID software and nicely described in the methods section of the GeneID publication.<\/p>\n<p>\u5728\u8fd0\u884c\u8fc7PASA\u8fc7\u540e\uff0c\u5219\u53ef\u4ee5\u8fdb\u884cORFs\u7684\u63d0\u53d6\uff1a<\/p>\n<pre>$ export PERL5LIB $PERL5LIB:$PASAHOME\/PerlLib\/\r\n$ $PASAHOME\/scripts\/pasa_asmbls_to_training_set.dbi -M \"$mysql_db_name:$mysql_server_name\" -p \"$user:$password\" -g genome_sample.fasta<\/pre>\n<p>\u8fd0\u884c\u7ed3\u679c\uff1a<br \/>\n<span style=\"color: #888888;\"><em>trainingSetCandidates.cds,.pep,.gff3<\/em> <\/span>:corresponds to all longest ORFs of minimum length extracted from the PASA alignment assemblies.<\/p>\n<p><span style=\"color: #888888;\"><em>trainingSetCandidates.cds.top_500_longest<\/em><\/span> :corresponds to the 500 top longest ORFs, thought to be most likely to represent correct protein-coding genes.<br \/>\n<em><br \/>\n<span style=\"color: #888888;\">trainingSetCandidates.cds.scores<\/span><\/em> :scores for all ORFs based on the Markov model.<\/p>\n<p><span style=\"color: #888888;\"><em>best_candidates.cds,.pep,.gff3,.bed<\/em> <\/span>:the best candidates based on their coding scores compared to random sequences of the same nucleotide composition.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>\u524d\u6587\u8bb2\u8ff0\u4e86\uff1aPASA\u7684\u5b89\u88c5\uff0c\u914d\u7f6e\u4e0e\u4e3b\u7a0b\u5e8f\u4f7f\u7528\u53c2\u6570\u3002\u672c\u6587\u5219\u91cd\u70b9\u8bb2\u8ff0PASA\u7684\u51e0\u79cd\u5e73 &hellip; <a href=\"http:\/\/www.chenlianfu.com\/?p=1181\">\u7ee7\u7eed\u9605\u8bfb <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":[],"categories":[],"tags":[],"_links":{"self":[{"href":"http:\/\/www.chenlianfu.com\/index.php?rest_route=\/wp\/v2\/posts\/1181"}],"collection":[{"href":"http:\/\/www.chenlianfu.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"http:\/\/www.chenlianfu.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"http:\/\/www.chenlianfu.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"http:\/\/www.chenlianfu.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=1181"}],"version-history":[{"count":24,"href":"http:\/\/www.chenlianfu.com\/index.php?rest_route=\/wp\/v2\/posts\/1181\/revisions"}],"predecessor-version":[{"id":1184,"href":"http:\/\/www.chenlianfu.com\/index.php?rest_route=\/wp\/v2\/posts\/1181\/revisions\/1184"}],"wp:attachment":[{"href":"http:\/\/www.chenlianfu.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=1181"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"http:\/\/www.chenlianfu.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=1181"},{"taxonomy":"post_tag","embeddable":true,"href":"http:\/\/www.chenlianfu.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=1181"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}