Well, my paper won't be submitted for a few months. But these issues must be dealt with sooner or later… it kind of comes with the territory of having users, which is the goal. I would say I don't mean to put undue urgency on the issue… but actually, I think I do mean to do it. Nothing focuses our current "solidification" effort better than actually trying to do something for real. "One big question that we need to answer at the policy level is, are we going to be a data repository?" Well…. let's contemplate this a moment. What I'm asking to do is publish a paper about an analysis I did in KBase. And you seem to be asking if we should even hold onto the data from the analysis. If this is where we're at right now… then we are not only dead in the water, but we're at the bottom of the ocean. How is this not completely a "no-brainer"? What am I missing here? What's with all the talk about provenance, etc if this isn't the goal? What is so magical about NCBI? I guarantee you our annotations are more consistent (and better) than theirs… and we deliver genome sequences just as well as they do. "There are probably good reasons NCBI makes you jump through so many hoops." Well, I think the reason is mostly because "they can". Perhaps I'm being overly cynical... "We need to settle on the data models and ids." Yeah sure. I can go along with this. Let's settle the ID issue right now. Genome IDs: "kb|g.####". I'm afraid the ship sailed long ago on that one, and I don't like the "|" any more than you do. Phenotypes: "kb|g.###.phenos.###". Speak now, or forever hold your peace. Data models. Well, I just don't see us setting those in stone right this minute, or anytime in the next 1-2 years. How about we decide what is most critical to have in the data model and make sure it's there? On Apr 17, 2013, at 1:07 PM, Paramvir Dehal <[email protected]> wrote:
One big question that we need to answer at the policy level is, are we going to be a data repository?