Posts

Showing posts with the label databases

Sharing Data, Information and Knowledge

LNCS 5071: Sharing Data, Information and Knowledge (25th British National Conference on Databases, BNCOD 25, Cardiff, UK, July 7-10, 2008). Some interesting papers: The Power of Data Visualisation to Aid Biodiversity Studies through Accurate Taxonomic Reconciliation Distributed Systems and Automated Biodiversity Informatics: Genomic Analysis and Geographic Visualization of Disease Evolution. Role Based Access to Support Collaboration in Healthcare Finding Data Resources in a Virtual Observatory Using SKOS Vocabularies Semantic Matching for the Medical Domain The Hyperdatabase Project – From the Vision to Realizations

Scientific and Statistical Database Management

Volume 5069 of LNCS titled " Scientific and Statistical Database Management " (20th International Conference, SSDBM 2008, Hong Kong, China, July 9-11, 2008) is just out. Some interesting papers: New Challenges in Petascale Scientific Databases Query Planning for Searching Inter-dependent Deep-Web Databases A Probabilistic Framework for Building Privacy-Preserving Synopses of Multi-dimensional Data ViP: A User-Centric View-Based Annotation Framework for Scientific Data Flexible Scientific Workflow Modeling Using Frames, Templates, and Dynamic Embedding Examining Statistics of Workflow Evolution Provenance: A First Study Adventures in the Blogosphere Ontology Database: A New Method for Semantic Modeling and an Application to Brainwave Data NB: I would be using DOIs for these articles as they are available but they do not seem to be registered yet with doi.org...

Extremely Large Databases

The First Workshop on Extremely Large Databases was held at the Stanford Linear Accelerator Center , October 2007. Many of the heavy hitters were there (Google, Yahoo, Microsoft, IBM, Oracle, Terrasoft, SLAC, NCSA, eBay, AT&T, etc) from industry, academia and science (? their classification). A report is available and I thought I'd touch on some of the more interesting things I found in it: Scale : Most have systems with > 100TB of data, with 20% of scientific databases > 1PB of data; All from industry reps had >100PB of data, with all having at least one system with >1PB Industry had single tables with > 1 trillion rows; science ~100 times smaller. Need for multi-trillion-row tables in Peak ingest: 1B rows per hour; 1B rows per day common " All users said that even though their databases were already growing rapidly, they would store even more data in databases if it were affordable. Estimates of the potential ranged from ten to one hundred times current...