Posts

Showing posts with the label scientific data

Job ad: Scientific Data Management Specialist

The following excerpt from an ad for Scientific Data Management Specialist suggests it bodes well for the prospects of this (relatively) nascent profession: Processing, soliciting, and providing assistance with data submissions for scientific data from genome sequencing and genotyping experiments into existing databases , analysis pipelines and associated data flows. Developing and improving the infrastructure supporting these systems. Required Skills Formal Education PhD Scripting experience in perl or related language Experience with SQL Experience with LINUX/UNIX Ability to use Microsoft Excel and related applications Proven record solving related problems Desired Skills Knowledge of genetics, especially human genetics Experience with large data sets XML/XSLT and related web based tools Experience with array data, especially expression or genotyping data C/C++ Experience with grid computing (LSF,SunGrid, etc.) QA filtering of genotype data (HWE, non-Mendelian segregation)...

Research Data Archiving = Volumes of data? Conference...

One of the implications of research data archiving is that there will likely be large datasets on a variety of subjects. How third parties will analyse these data in a scalable fashion has not entirely been addressed. But the recent conference examines some of the issues for a class of data, images and signals : Advances in Mass Data Analysis of Images and Signals in Medicine, Biotechnology, Chemistry and Food Industry (Third International Conference, MDA 2008 Leipzig, Germany, July 14, 2008) and includes: Burcu Yılmaz, Mehmet Göktürk, Natalie Shvets (2008). User Assisted Substructure Extraction in Molecular Data Mining. Advances in Mass Data Analysis of Images and Signals in Medicine, Biotechnology, Chemistry and Food Industry, 5108 , 12-26 DOI: 10.1007/978-3-540-70715-8_4 Franco Chiarugi, Sara Colantonio, Dimitra Emmanouilidou, Davide Moroni, Ovidio Salvetti (2008). Biomedical Signal and Image Processing for Decision Support in Heart Failure. Advances in Mass Data Analysis of Imag...

Open Access to calculated chemical properties of 12M substances

ChemStar (Computed Molecular Data) is an Open Access database of the calculated chemical properties of over 12M substances from PubChem. The process - which demands significant computing power - was done is a distributed fashion using JavaRMI and is described in the paper below. The properties are calculated using three different methods: JChem , JOELIB and MOE . The site allows the perusal of the properties, with a Marvin -based applet for viewing structures. The Java sourcecode for the distributed computing is available . Karthikeyan, M., Krishnan, S., Pandey, A.K., Bender, A., Tropsha, A. (2008). Distributed Chemical Computing Using ChemStar: An Open Source Java Remote Method Invocation Architecture Applied to Large Scale Molecular Data from PubChem. Journal of Chemical Information and Modeling, 48 (4), 691-703. DOI: 10.1021/ci700334f

Scientific and Statistical Database Management

Volume 5069 of LNCS titled " Scientific and Statistical Database Management " (20th International Conference, SSDBM 2008, Hong Kong, China, July 9-11, 2008) is just out. Some interesting papers: New Challenges in Petascale Scientific Databases Query Planning for Searching Inter-dependent Deep-Web Databases A Probabilistic Framework for Building Privacy-Preserving Synopses of Multi-dimensional Data ViP: A User-Centric View-Based Annotation Framework for Scientific Data Flexible Scientific Workflow Modeling Using Frames, Templates, and Dynamic Embedding Examining Statistics of Workflow Evolution Provenance: A First Study Adventures in the Blogosphere Ontology Database: A New Method for Semantic Modeling and an Application to Brainwave Data NB: I would be using DOIs for these articles as they are available but they do not seem to be registered yet with doi.org...

FREE THE ARTICLES! (Full-text for researchers & scientists and their machines)

At a recent plenary I gave [ earlier post ] at the Colorado Association of Research Libraries Next Gen Library Interfaces conference, I went a little off-script and was educating (/haranguing) the mostly librarian audience about the present-and-near-future importance of the accessibility of full-text research articles to their researchers and scientists. By accessibility of full-text I didn't mean the ability of a human to access the PDF or HTML of an article via a web browser: I was referring to the machine- accessibility of the text contained in the article (and the metadata and the citation information). I was concerned because of the increasing number of discipline-specific tools that use full-text (& metadata & citations) to allow users (via text mining, semantic analysis, etc.) to navigate, analyze and discover new ideas and relationships, from the research literature. The general label for this kind of research is ' literature-based discovery ', where new ...

AAAS Meeting: "Managing and Preserving Scientific Data: Emerging Perspectives on a Global Basis"

I will be participating in this year's AAAS meeting in Boston with a presentation entitled " Canadian Initiative To Develop a National Strategy " at the session "Managing and Preserving Scientific Data: Emerging Perspectives on a Global Basis", moderated by Bonnie Carrol. The other presentations: U.S. National Initiatives: Strategic Plan for Scientific Data Management and Preservation. Christopher L. Greer, National Science Foundation, USA U.K. Initiatives and Perspectives on Managing and Preserving Scientific Data. Liz Lyon, University of Bath, UK European Framework: Promote Access and Preserve Research Results for Future Generations. Carlos Morais-Pires, European Commission

Proposed Standard for Citing Quantitative Research Data

I must have been asleep this spring to have missed this interesting article, A Proposed Standard for the Scholarly Citation of Quantitative Data (M. Altman & G. King, D-Lib Magazine , March/April 2007, Volume 13 Number 3/4) which introduces a reasonable specification allowing the citation of quantitative data. This is a needed specification, which will hopefully increase the real (and perceived) value of datasets to researchers, the people who evaluate them and the people who fund them. This allows datasets to be a measurable metric by which a researcher's performance can be measured: through usage and peer-review. " We propose that citations to numerical data include, at a minimum, six required components. The first three components are traditional, directly paralleling print documents. ... Thus, we add three components using modern technology, each of which is designed to persist even when the technology changes: a unique global identifier, a universal numeric fingerpri...

CSIRO research program takes data leadership

CSIRO has announced a new program, Terabyte Science , (press release: From molecules to the Milky Way: dealing with the data deluge ) that is oriented around dealing with the issues of the large volumes of data generated by much of modern science. While this includes the difficult problem of the management of large volumes of data, this program will also focus on new ways to analyse and exploit this data. Kudos to CSIRO for recognizing this issue and realizing an important program for data science, both to their own country and the rest of us. Hopefully other such initiatives will take hold in other countries. And perhaps we will be seeing some of their work published in CODATA 's Data Science Journal . (Disclosure: I am an observer on the Canadian National Committee for CODATA ).

IJDL Special Issue: Connecting digital libraries to eScience

The International Journal on Digital Libraries has a special issue entitled " Connecting digital libraries to eScience". I haven't had a chance to read any of the articles, but they look very interesting, and include some discussion on various scientific data issues, collaboration, repositories, research infrastructure, etc: Connecting digital libraries to eScience: the future of scientific scholarship . Michael Wright, Tamara Sumner, Reagan Moore, Traugott Koch Not by metadata alone: the use of diverse forms of knowledge to locate data for reuse . Ann Zimmerman Little science confronts the data deluge: habitat ecology, embedded sensor networks, and digital libraries . Christine L. Borgman, Jillian C. Wallis, Noel Enyedy Collaborative eScience libraries . Linn Marks Collins, Mark L. B. Martinez, Ketan K. Mane, James E. Powell, Chad M. Kieffer, Tiago Simas, Susan K. Heckethorn, Kathryn R. Varjabedian, Miriam E. Blake, Richard E. Luce Pathways: augmenting interoperabilit...

New JISC Data Sharing Documents

As part of its DISC -UK DataShare project , JISC has released two documents: DISC-UK DataShare: State-of-the-Art Review , Harry Gibbs Data Sharing Continuum graphic , Robin Rice The former is a summary of recent projects and policy, and introduced me to a number of projects and initiatives that I hadn't previously known about. The latter is a well thought-out view of the data sharing continuum, showing us where we have been (and perhaps for some of us, still are!) and a good idea of where we will/should be going. A good graphic to show to a manager trying to understand the big picture.

Sustainable Digital Data Preservation and Access Network Partners (DATANET) CFP

The US National Science Foundation Office of Cyberinfrastructure has a call for proposals . From the call: The new types of organizations envisioned in this solicitation will integrate library and archival sciences, cyberinfrastructure, computer and information sciences, and domain science expertise to: provide reliable digital preservation, access, integration, and analysis capabilities for science and/or engineering data over a decades-long timeline; continuously anticipate and adapt to changes in technologies and in user needs and expectations; engage at the frontiers of computer and information science and cyberinfrastructure with research and development to drive the leading edge forward; and serve as component elements of an interoperable data preservation and access network. ...these exemplar organizations can serve as the basis for rational investment in digital preservation and access by diverse sectors of society at the local, regional, national, and international levels, ...

New Zealand Science and Open Access

In " An Information Revolution ", David Penman discusses Open Access and Open Data (especially as applied to government-funded research) in general, and more specifically as applied to New Zealand science and scientists. While there is some good news: "The Foundation for Research, Science and Technology is now reviewing its data policy and moving towards the norm for the OECD – greater open access for publicly-funded data. Rather than the research provider deciding on access, all information is openly and freely available unless restrictions such as national security, environmental damage (eg, the GPS co-ordinates of threatened species), or clear commercial disadvantage can be justified." He has some blunt - and appropriate - words for NZ scientists: Our researchers will also have to change. No longer can they sit with filing cabinets full of data waiting for the definitive experiment or the life time monograph. Publish quickly in electronic media, make your data a...

Australia talks about Research data archiving

I see how the Australians appear to have the good fortune of having the discussion on research data archiving moving forward, as suggested by the upcoming meeting in September, " Long-Lived Collections: the Future of Australia's research data " at the National Library of Australia. This meeting is a follow-up to some very good efforts, including the Australian government's " Data for Science (DFS)" prepared for the Prime Minister’s Science, Engineering and Innovation Council, and the Australian Partnership for Sustainable Repositories' " Sustainability Issues for Australian Research Data: The report of the Australian eResearch Sustainability Survey Project ". I can only be envious of this activity, given the -- unfortunately -- almost complete vacuum of activity following the release of Canada's National Consultation on Access to Scientific Research Data (NCASRD). The two reports - DfS and NCASRD - are very similar in scope and in reco...

Data Archiving of Publicly Funded Research in Canada

Carol Perry presented this revealing study at last year's Access & Privacy Workshop 2006 held in Toronto. Its objectives were: "To assess the attitudes of academic researchers regarding the archiving of data resulting from publicly funded research To assess impediments to the creation of a national data archive program in Canada" She randomly polled 173 SSHRC grant recipients for 2004-2005 (with 75 respondents). Her results: "41% indicated they had current plans to archive their research data Of these, only 18.7% identified an established data archive as a deposit site for their data. 72% were not aware of SSHRC’s mandatory data archiving policy for all grant recipients 90% were not aware that Canada is a recent signatory to the OECD declaration on access to publicly funded data ." and •" In 2001: 60% favoured a national data archive 39% analyzed data created by others •In 2006: 69% favoured a national data archive 48% analyzed data created by others...

"Scaleable Knowledge Discovery through Grid Workflows"

USC 's Information Sciences Institute has been awarded $13.8M for this project , which aims to solve the massive data glut that swamps many disciplines by automating scientific workflows . Using grid computing and other advanced technologies (such as A.I. and Semantic Web ), a workflow architecture allows for the capture of complex scientidic workflows and their application in a distributed fashion on data sets in an efficient manner. This is a very ambitious (and well funded) project. Some of the (even more) interesting parts of this project include (from the article): "...investigate mechanisms to support autonomous and robust execution of concurrent workflows over continuously changing data...learning techniques to improve the performance of the workflow system by exploiting an episodic memory of prior workflow executions." episodic memory (Wikipedia)

EU: 50M Euro for digital repositories for scientific data

EU: 50M Euro for digital repositories for scientific data The European Commissioner for Science and Research has announced ( Communication from the Commission to the European Parliament, the Council and the European Economic and Social Committee on Scientific Information in the Digital Age: Access, Dissemination and Preservation ) 50M euros to be put aside for a digital repository for scientific data. Wow! This is amazing news. I am hoping that this might spur Canada to wake up and do something similar, especially after having two reports discussing similar activities in Canada ( National Consultation on Access to Scientific Research Data and National Data Archive Consultation Final Report: Building Infrastructure for Access to and Preservation of Research Data in Canada )