Posts

Showing posts with the label e-science

ARL Report: E-Science and Data Support Services

The U.S. Association of Research Libraries (ARL) has produced a new report ( E-Science and Data Support Services ).

Microsoft Research Faculty Summit 2008: Publishing and Research Tools for Academics

The Microsoft Research Faculty Summit 2008 included Publishing and Research Tools for Academics that allows for archival annotation and structuring of Microsoft software produced documents, as well as supporting PubMed Central format . Additional sessions of interest: The Cyberspace Connection – Impact on Individuals, Society, and Research New Developments in Scholarly Communication Reflections on Directions in Artificial Intelligence What Will Be the Impact of Cloud Services on Science? Spotlights on Interdisciplinary Artificial Intelligence Research AI, Sensing, and Optimized Information Gathering: Trends and Directions Ontological Myths: Reducing the Confusion Social Networking and Semantics Toward Situated Interaction Statistical Machine Translation Research at Microsoft Research Interactive Machine Learning: Challenges, Methods, and Applications Information Extraction from Documents and Queries Contexts in Computer Science Education The Future of Research Clouds REAssess: Resou...

Scientific and Statistical Database Management

Volume 5069 of LNCS titled " Scientific and Statistical Database Management " (20th International Conference, SSDBM 2008, Hong Kong, China, July 9-11, 2008) is just out. Some interesting papers: New Challenges in Petascale Scientific Databases Query Planning for Searching Inter-dependent Deep-Web Databases A Probabilistic Framework for Building Privacy-Preserving Synopses of Multi-dimensional Data ViP: A User-Centric View-Based Annotation Framework for Scientific Data Flexible Scientific Workflow Modeling Using Frames, Templates, and Dynamic Embedding Examining Statistics of Workflow Evolution Provenance: A First Study Adventures in the Blogosphere Ontology Database: A New Method for Semantic Modeling and an Application to Brainwave Data NB: I would be using DOIs for these articles as they are available but they do not seem to be registered yet with doi.org...

"Libraries in the Converging Worlds of Open Data, E-Research, and Web 2.0"

This looks like an interesting article ( Libraries in the Converging Worlds of Open Data, E-Research, and Web 2.0 , Stuart MacDonald, March/April issue of ONLINE magazine) but I don't have a subscription so I can't really comment much on it. Ironic that the abstract mentions Peter Suber's Open Access News blog.... ;-) Abstract: " The new forms of research enabled by the latest technologies bring about collaboration among researchers in different locations, institutions, and even disciplines. These new collaborations have two key features -- the prodigious use and production of data. This data-centric research manifests itself in such concepts as e-science, cyberinfrastructure, or e-research. Over the last decade there has been much discussion about the merits of open standards, open source software, open access to scholarly publications, and most recently open data. There are a range of authoritative weblogs that address the open movement, some of which include: 1. DC...

Hadoop + EC2 + S3 = Super alternatives for researchers (& real people too!)

I recently discovered and have been inspired by a real-world and non-trivial (in space and in time) application of Hadoop (Open Source implementation of Google's MapReduce ) combined with the Amazon Simple Storage Service ( Amazon S3 ) and the Amazon Elastic Compute Cloud ( Amazon EC2 ). The project was to convert pre-1922 New York Times articles-as-scanned-TIFF-images into PDFs of the articles: Recipe: 4 TB of data loaded to S3 (TIFF images) + Hadoop (+ Java Advanced Imaging and various glue) + 100 EC2 instances + 24 hours = 11 million PDFs , 1.5 TB on S3 Unfortunately, the developer ( Derek Gottfrid ) did not say how much this cost the NYT. But here is my back-of-the-envelope calculation (using the Amazon S3/EC2 FAQ ): EC2: $0.10 per instance-hour x 100 instances x 24hrs = $240 S3 : $0.15 per GB-Month x 4500 GB x ~1.5/31 months = ~$33 + $0.10 per GB of data transferred in x 4000 GB = $400 + $0.13 per GB of data transferred out x 1500 GB = $195 Total: = ~$868 Not unre...

Plant Science "Grand Challenges" Cyberinfrastructure Funded by NSF

The NSF announced today announced it would be awarding $50M to the iPlant Collective , a " dynamic web portal for the Plant Science Cyberinfrastructure Collaborative Community " to address the ' grand challenges' in plant science. The effort will create a global centre bringing together (virtually and actually) computer scientists, information scientists and plant scientists to work on projects untenable until this project due to such issues as complexity, scale, discipline boundaries, lack of collaboration structures, etc. The centre "... will bring together and leverage the resources and information generated through the National Plant Genome Initiative , enabling more breadth and depth of research in every aspect of plant science " and will serve as a model for other disciplines on how collaborative cyberinfrastructure can be applied.

AAAS Meeting: "Managing and Preserving Scientific Data: Emerging Perspectives on a Global Basis"

I will be participating in this year's AAAS meeting in Boston with a presentation entitled " Canadian Initiative To Develop a National Strategy " at the session "Managing and Preserving Scientific Data: Emerging Perspectives on a Global Basis", moderated by Bonnie Carrol. The other presentations: U.S. National Initiatives: Strategic Plan for Scientific Data Management and Preservation. Christopher L. Greer, National Science Foundation, USA U.K. Initiatives and Perspectives on Managing and Preserving Scientific Data. Liz Lyon, University of Bath, UK European Framework: Promote Access and Preserve Research Results for Future Generations. Carlos Morais-Pires, European Commission

Proposed Standard for Citing Quantitative Research Data

I must have been asleep this spring to have missed this interesting article, A Proposed Standard for the Scholarly Citation of Quantitative Data (M. Altman & G. King, D-Lib Magazine , March/April 2007, Volume 13 Number 3/4) which introduces a reasonable specification allowing the citation of quantitative data. This is a needed specification, which will hopefully increase the real (and perceived) value of datasets to researchers, the people who evaluate them and the people who fund them. This allows datasets to be a measurable metric by which a researcher's performance can be measured: through usage and peer-review. " We propose that citations to numerical data include, at a minimum, six required components. The first three components are traditional, directly paralleling print documents. ... Thus, we add three components using modern technology, each of which is designed to persist even when the technology changes: a unique global identifier, a universal numeric fingerpri...

"ERC Scientific Council Guidelines for Open Access"

In my rather hectic December I missed this publication on Dec 17 2007 of the ERC Scientific Council Guidelines for Open Access. From the document's " interim position on open access ": 1. The ERC requires that all peer-reviewed publications from ERC-funded research projects be deposited on publication into an appropriate research repository where available, such as PubMed Central, ArXiv or an institutional repository, and subsequently made Open Access within 6 months of publication . [Emphasis added] 2. The ERC considers essential that primary data - which in the life sciences for example could comprise data such as nucleotide/protein sequences, macromolecular atomic coordinates and anonymized epidemiological data - are deposited to the relevant databases as soon as possible, preferably immediately after publication and in any case not later than 6 months after the date of publication. [Emphasis added] Thanks to Mary Zborowski, CISTI, NRC and CODATA for pointing this out...

ARL Report released: Library Support for E-science

Association of Research Libraries (ARL) has just released the report " Agenda for Developing E-Science in Research Libraries ". Among other activities, this report specifically discusses the Canadian context, including the National Consultation on Access to Scientific Research Data (NCASRD) and the recent (2007) Library and Archives Canada's release draft of its Canadian Digital Information Strategy . Agenda for Developing E-Science in Research Libraries table of contents: E-Science: Implications for Research Practice National and International Context for E-Science Critical Areas for Research Library Engagement Data Issues and New Genres of Scholarly Communication Virtual Organizations Policy Development Current Library Capability to Support E-Science E-Science Task Force Recommendations: Outcomes, Strategies, Actions Structure and Process for ARL Agenda Develop Knowledgeable Community Develop Skilled Workforce Contribute to Research Infrastructure Develop Policy From ...

CSIRO research program takes data leadership

CSIRO has announced a new program, Terabyte Science , (press release: From molecules to the Milky Way: dealing with the data deluge ) that is oriented around dealing with the issues of the large volumes of data generated by much of modern science. While this includes the difficult problem of the management of large volumes of data, this program will also focus on new ways to analyse and exploit this data. Kudos to CSIRO for recognizing this issue and realizing an important program for data science, both to their own country and the rest of us. Hopefully other such initiatives will take hold in other countries. And perhaps we will be seeing some of their work published in CODATA 's Data Science Journal . (Disclosure: I am an observer on the Canadian National Committee for CODATA ).

IJDL Special Issue: Connecting digital libraries to eScience

The International Journal on Digital Libraries has a special issue entitled " Connecting digital libraries to eScience". I haven't had a chance to read any of the articles, but they look very interesting, and include some discussion on various scientific data issues, collaboration, repositories, research infrastructure, etc: Connecting digital libraries to eScience: the future of scientific scholarship . Michael Wright, Tamara Sumner, Reagan Moore, Traugott Koch Not by metadata alone: the use of diverse forms of knowledge to locate data for reuse . Ann Zimmerman Little science confronts the data deluge: habitat ecology, embedded sensor networks, and digital libraries . Christine L. Borgman, Jillian C. Wallis, Noel Enyedy Collaborative eScience libraries . Linn Marks Collins, Mark L. B. Martinez, Ketan K. Mane, James E. Powell, Chad M. Kieffer, Tiago Simas, Susan K. Heckethorn, Kathryn R. Varjabedian, Miriam E. Blake, Richard E. Luce Pathways: augmenting interoperabilit...

Sustainable Digital Data Preservation and Access Network Partners (DATANET) CFP

The US National Science Foundation Office of Cyberinfrastructure has a call for proposals . From the call: The new types of organizations envisioned in this solicitation will integrate library and archival sciences, cyberinfrastructure, computer and information sciences, and domain science expertise to: provide reliable digital preservation, access, integration, and analysis capabilities for science and/or engineering data over a decades-long timeline; continuously anticipate and adapt to changes in technologies and in user needs and expectations; engage at the frontiers of computer and information science and cyberinfrastructure with research and development to drive the leading edge forward; and serve as component elements of an interoperable data preservation and access network. ...these exemplar organizations can serve as the basis for rational investment in digital preservation and access by diverse sectors of society at the local, regional, national, and international levels, ...

Fedora to grow to include open access publishing, eScience, and eScholarship

Sandy Payette has plans to expand Fedora to support open access publishing, eScience and eScholarship. With a recent $4.9M grant from the Moore foundation, it looks like she might have the opportunity to do this...

Clifford Lynch on Cyberinfrastructure and E-Research

Clifford Lynch 's closing keynote to the 2007 Seminars On Academic Computing entitled " The Institutional Challenges of Cyberinfrastructure and E-Research " is now available as a podcast . Abstract: It has become clear that scholarly practice and scholarly communication across a wide range of disciplines are being transfigured by a series of developments in IT and networked information. While this has been widely discussed at the national and international levels in the context of large-scale advanced scientific projects, the challenges at the level of individual universities and colleges may prove more complex and more difficult. This presentation will focus on these challenges, as well as the development of truly institution-wide strategies that can support and advance the promises of e-research.

"Scaleable Knowledge Discovery through Grid Workflows"

USC 's Information Sciences Institute has been awarded $13.8M for this project , which aims to solve the massive data glut that swamps many disciplines by automating scientific workflows . Using grid computing and other advanced technologies (such as A.I. and Semantic Web ), a workflow architecture allows for the capture of complex scientidic workflows and their application in a distributed fashion on data sets in an efficient manner. This is a very ambitious (and well funded) project. Some of the (even more) interesting parts of this project include (from the article): "...investigate mechanisms to support autonomous and robust execution of concurrent workflows over continuously changing data...learning techniques to improve the performance of the workflow system by exploiting an episodic memory of prior workflow executions." episodic memory (Wikipedia)
AAAI2007's workshop on Semantic e-Science CFP deadline extended I see that the CFP for this workshop at AAAI 2007 (in Vancouver, July 22-26) has been extended until April 12 2007.