Posts

Showing posts with the label text mining

The (near) Future of Research Articles

Image
Rod Page 's demo for his Elsevier Grand Challenge submission (" Towards realising Darwin’s dream: setting the trees free ") shows the type of enrichment of biological - if not all research - articles that is quickly becoming possible. Taking a published article (" Mitochondrial paraphyly in a polymorphic poison frog species (Dendrobatidae; D. pumilio "), various additional biological, geographical and other metadata are extracted and added to a web page for the article. These include: Map showing all localities mentioned in the paper, with their enclosing polygon List of other studies which have samples in area enclosed by the study polygon Each of the following are linked through to their underlying databases (such as NIH accession number and NCBI nucleotide viewer or linked to ubio taxonomic name viewer record: List of sequence features (such as genes) in the article List of taxa sequenced in the article List of gene sequences cited by the article An image c...

New Open Access Criterion: Support access by machines (m2m)

Related to my last posting ( FREE THE ARTICLES! (full-text for researchers & scientists and their machines) ) and in the light of Peter Murray-Rust's recent annoying discovery that he cannot text-mine Pubmed Central ( Can I data- and Text-mine Pubmed Central? ), I would like to suggest an additional criterion to the definition of Open Access: Open Access must include access by machines : At minimum one must allow crawls of the site/content or (to reduce the impact of badly configured crawlers) create a compressed XML file containing all metadata and either content, or direct links to content and make it available for download (and if bandwidth is still an issue put it on a P2P network like BitTorrent ). Preferable is to offer some kind of API (OTMI) or protocol (OAI-PMH) to get at content and metadata and citations. Better is to offer access to the XML of the articles in addition to the PDF and/or HTML; if the XML actually has some semantic content, then we are approaching th...

FREE THE ARTICLES! (Full-text for researchers & scientists and their machines)

At a recent plenary I gave [ earlier post ] at the Colorado Association of Research Libraries Next Gen Library Interfaces conference, I went a little off-script and was educating (/haranguing) the mostly librarian audience about the present-and-near-future importance of the accessibility of full-text research articles to their researchers and scientists. By accessibility of full-text I didn't mean the ability of a human to access the PDF or HTML of an article via a web browser: I was referring to the machine- accessibility of the text contained in the article (and the metadata and the citation information). I was concerned because of the increasing number of discipline-specific tools that use full-text (& metadata & citations) to allow users (via text mining, semantic analysis, etc.) to navigate, analyze and discover new ideas and relationships, from the research literature. The general label for this kind of research is ' literature-based discovery ', where new ...
Tapping the power of text mining In his closing plenary to the Access 2006 conference in Ottawa, Clifford Lynch listed text mining as one of the exciting areas of activity for the near future, soon (hopefully!) realizing its potential for discovery on large text corpora. In the September 2006 issue of Communications of the ACM, Fan et al. have a good general introduction to this area. Fan, W., Wallace, L., Rich, S., and Zhang, Z. 2006. Tapping the power of text mining. Commun. ACM 49, 9 (Sep. 2006), 76-82. DOI= http://doi.acm.org/10.1145/1151030.1151032 More text mining Wikipedia text-mining.org New Zealand Digital Library