On Mar 14, 2013, at 2:01 PM, Folker Meyer <[email protected]> wrote:
- Many users find metadata annoying and don't like to provide it. Our approach has been to tell them they get better service if they provide the metadata and also provide tools that make minimal metadata easy --> Metazen. - Marcin is right that richer metadata is important (this is solved via the "environmental packages") and the notion of different levels of metadata. Richness and ease of use are somewhat opposite goals, Doreen points out that we will have to coerce our users to do the right thing. One way might be to have tools in KBase that require certain sets of metadata to work e.g. the use of certain CVs for experimental conditions. - We probably should go for multiple levels of metadata completeness, encouraging the users to add as much as they can, while requiring a minimal subset. At least that is what we did in MG-RAST.
Hope I'm not going too off-topic here. Sunita and I have been brainstorming just a few days ago about building functionality for rudimentary annotation of ontology terms from cited literature, which is very often (always?) associated with GEO data sets. This service would essentially use natural language processing[1] to mine papers for these mappings. This would be in the very common case that GEO data is not already associated with metadata. An NLP approach, as far as current pipelines go, have very high false-positive rates. But we can model it in such way that captures a measure of confidence about the mapping, so unreliable mappings aren't taken as Truth. Maybe these annotations can serve as a very liberal baseline for categorizing entities, when users don't care about providing the mappings themselves. Shiran [1] Disclaimer: I'm NOT an NLP expert! :)