Posts

Showing posts with the label LuSql

University visitor @ Australian National University

Image
Tomorrow is (sadly) my last official day * as a university visitor at the Australian National University (ANU), Canberra . I've been here since late June, invited by ANU adjunct and Funnelback chief scientist (and ex-CSIRO ) David Hawking , to visit the Algorithms and Data Research Group , School of Computer Science , College of Engineering and Computer Science . I was installed in a lovely office looking out into the campus, where I've been working on large scale journal visualization, a continuation of the Torngat project . I've been working on a couple of things, including applying Mulan to the multi-label problem of the corpus I am working with, so I can get precision and recall to evaluate this method empirically. My productivity has been hampered by a recurring stomach problem (which appears to be gone this last week: yay!), so I've not progressed as much as I would have wanted to.... :-( At the end of last week I gave a presentation at CSIRO (in the same bui...

Presenting at Code4Lib-North

Using Open Source Tools for Visualization and Semantic Mapping in a Large Scale Article Digital Library [slideshare] - Questions: Anyone working in the cloud? Converting 4TB of TIFFS to PDF in 24hrs: Hadoop + EC2 + S3 = Super alternatives for researchers (& real people too!)

Project Torngat: Building Large-Scale Semantic 'Maps of Science' with LuSql, Lucene, Semantic Vectors, R and Processing from Full-Text

Image
Project Torngat is a research project here at NRC - CISTI   [ Note that I am no longer at CISTI and that I am now continuing this work at Carleton University - GN 2010 04 07 ] that looks to use the full-text of journal articles to construct semantic journal maps for use in -- among other things -- projecting article search results onto the map to visualize the results and support interactive exploration and discovery of related articles, term and journals. Starting with 5.7 million full-text articles from 2200+ journals (mostly science, technology and medical (STM)), and using LuSql , Lucene , Semantic Vectors , R , and processing , a two dimensional mapping of a 512 dimension semantic space was created which revealed an excellent correspondence with the 23 human-created journal categories: Semantic Journal Space of 2231 Journals Scaled to Two Dimensions This initial work was initiated to find a technique that would scale, and follow-up work is looking at integrati...

code4lib update: LuSql talk done; Lucene, Solr links

Gave my LuSql talk today at code4lib2009 and didn't get cut down by any Solr/Lucene dudes! Met Erik Hatcher of Lucene/Solr fame (and now of Lucid Imagination fame) & hopefully we can collaborate on some Lucene/indexing Solr stuff in the future. I also spoke with Tom Burton-West of UMich about Lucene indexing and search performance for their 1M+ Google Books index (they use Solr). These are documents that are a lot longer than the STM articles I work with. They have 220GB sized indexes and - as they have to keep stops words for their Humanities for phrase searching - suffer from poor query performance (despite 32GB RAM). I pointed to some of my previous work on high performance indexing and searching [ 1 , 2 , 3 ]. I'd like to get at their data to examine some performance issues in Lucene, both on the indexing and searching side. I was wondering if Solr is configurable for the initial/max number of IndexSearchers. I couldn't find this in the Solr wiki, but did see in...

code4lib pre-conference: Linked Data et al...

I am at the exciting and arcane code4lib 2009 conference here in Providence, Rhode Island. Right now at the pre-conference called LinkedData . on Linked Data . I had forgotten that Rhode Island and more specifically Providence, are the old stomping grounds (and location for many short stories and novels) of H.P. Lovecraft . And - this morning - I was talking to Ross Singer about this, and realised how this all made sense: when I first met Ross at an Access conference a number of years ago, the first thing I thought on meeting him was, " Chthulu "! He of course denied being one of the Elder Things and then levitated across the room from me. But I think this explains a lot of things... ;-) We will have to see what other Links I make at this conference. :-) Oh, BTW I will be giving a presentation tomorrow morning on LuSql . Feel free to drop in. :-)

Lucene 2.3.1 vs 2.4 benchmarks using LuSql

I have been doing some indexing performance tests with LuSql , and have some numbers comparing Lucene 2.3.1 with 2.4. Despite some discussion about 2.4 having poorer indexing performance, my tests with LuSql 0.9 suggest otherwise: Lucene 2.3.1 Number of records added= 2000000 Optimizing index Closing index Optimizing index time: 311 seconds Closing JDBC: result set Closing JDBC: statement Closing JDBC: connection *********** Elapsed time: 854 seconds 15m 18s Lucene 2.4 Number of records added= 2000000 Optimizing index Closing index Optimizing index time: 322 seconds Closing JDBC: result set Closing JDBC: statement Closing JDBC: connection *********** Elapsed time: 759 seconds 12m 39s Index size: 3.7GB. It is interesting that the overall indexing time is significantly less, but the optimizing time is slightly higher. Data, hardware and system configuration: as per my previous Lucene benchmarking . Note that this is a simple benchmark, so YMWV. This benchmark was done with the LuSql de...