Re: [Kbase-MicroComm] Starting a discussion of metadata in KBase
Hi Folker, Thanks for sending this - Yes, there are lots of thoughts on metadata. As you may know MicrobesOnline also stores metadata, and I venture to say that what we have for our in-house experiments and what Gavin has implemented in the Experiment service is a helpful standard/ example. Some pieces of that standard are: - Standardization across data sets, meaning that the terminology and content are consistent for different experiments. I would say this goes beyond CV in that the metadata itself is structured - perhaps this can be thought of as ordered metadata with required fields. In effect this means the data are machine readable and can be incorporated in automation. BTW this also includes strain sequence information. - Related to the above, the same people performed the experiments and encoded the metadata, which is rather luxurious. - The experiment compendium actually 'tests' different aspects of the conditions and this is reflected in the metadata. - The metadata is complete enough that one can build a virtual media for modeling etc. Much of this overlaps with your document. I know, part of the above is experimental design for which we usually have no control, but it does present a sort of best case. Regarding the last point for modeling, it may be hard to create virtual environments but perhaps also worth thinking about, I mentioned in house experiments and there the situation is nice. For the rest of the data, e.g. imports from GEO, things are not as nice. In fact they can be pretty horrible, up to and including lines of free text with no controlled vocabulary. It would be great to solve this problem in general but that may be a significant effort. Really that means devising a system to determine condition terms (CV) from free text and ensuring the term extraction is meaningful. In addition this is somewhat domain specific with say microarray conditions being mostly orthogonal to environmental data etc. Here the question is how your proposal applies to existing data in KBase or data that is automatically imported in batch. These will likely form the bulk of KBase data in the immediate future. Based on experience it will be hard to get users to upload their data through an interface, but not impossible - this depends on the 'reward'. Clearly having consistent meta data is enabling for integration and analysis. For example, knowing the similarity/ difference in conditions between say a metabolic and a phenotype experiment is rich information. If anything we should at least have the ability to 'match' conditions, or even score metadata similarity - perhaps this could be done with free text as well but without the need for CV exactly. I do like the idea of different levels of metadata quality but am worried that entropy will drive most of the cases to a less useful state. I believe the main question is - what does KBase plan to do with the metadata for different data types? Is it mostly labels for the UIs and text search? Personally I imagine being able to link different datasets by their conditions (based on a score) and performing enrichment analysis on the metadata terms ... Marcin On Mar 14, 2013, at 8:53 AM, Folker Meyer <[email protected]> wrote:
Hi all,
we have had several emails/discussions with several groups now on metadata in KBase. Now a couple of people are talking about using the API to store metadata.
There is a lot of prior discussion of metadata and metadata standards outside KBase and I think it is important to find a solution that integrates well with the external standards community. And in addition, rather than having a solution creep into place, I think it would be good to get in front of this and define what it is that we (as KBase) want.
As a result I mentioned in a recent discussion the fact that we have an implementation of this in MG-RAST that might be worth looking at. As a punishment for this I was asked to put together a document outlining our solution for this in MG-RAST (see attached).
There are some questions that I see us faced with:
- do we need a single metadata strategy for KBase - what aspects of metadata are required for each domain - what type of solution are people comfortable with - what are our customers comfortable with
I am sure I am missing tons of stuff that we need to think about and agree on a solution for, but hopefully this will be food for thought.
Best, Folker
<TheMG-RASTMetadatasolution.pdf> _______________________________________________ Kbase-microcomm mailing list [email protected] https://lists.kbase.us/mailman/listinfo/kbase-microcomm
participants (1)
-
Marcin