"If we are going to link in to workspaces then they need to be made immutable .. Which means there needs to be a transition from a user workspace to an archived reference workspace/data store. Once that transition happens its a reference bel object and "owned" by kbase not the user. It becomes a public dataset etc. Genomes if they are public should be in the CDM and treated the same as other public genomes. (Btw I'm discussing with jgi the issue of publishing genomes without going through NCBI. They wish to do this as well since they can no longer afford the overhead of iterating with NCBI on genome submission) .. Jim Bristow and I would like to do this in a coordinated way starting ona particular date." I agree with all of this. I like the idea of a special type of "Workspace" or "Project" that is a "Published Workspace" and is immutable. I would argue that the public genomes should be in the CDM AND the workspace/project space. I say this, because annotations and gene calls in the CDM may periodically change (I think that's the plan). So we still need the workspace version to serve as a permanent snapshot of the genomes as they stood when the analysis was first performed (I think). "You need something like the MIGS, MIM and MIBA (I made up the last two) standards. There are two aspects of this. The information we require for these datasets and the metadata mechanism. I strongly believe we need a single metadata mechanism for all of kbase that works for data in all of the stores (rather than one for each store or even worse a different mechanism for each type). Folker has proposed a mechanism that with some thought could be adopted as the scheme for all the stores. Each core data type needs to have a well defined set of items they want for each data set. I believe the mechanism proposed can support that." Yeah, I think Folker is way ahead of all of us on thinking about the metadata issue. I agree we should make every effort to adopt the system he's developed for this system "kbase-wide". I really like Tom's suggestion of a sprint team designated to the issue of a metadata and ID policy for KBase. I think his suggestion of people for the sprint team is also spot on: Chris, Folker, Paramvir, Annette, Shiran If someone else has an urgent burning desire to be on this team, they should volunteer. Then hopefully we can arrange some conference calls and prep a document with a firm KBase policy on the IDs and metadata plans, and identify the services that might be needed to support this? Does this seem amenable to everyone? Chris On Apr 18, 2013, at 8:12 AM, Rick stevens <[email protected]> wrote:
Sent from my iPad
On Apr 17, 2013, at 12:04 AM, Christopher Henry <[email protected]> wrote:
Hi all,
I started this thread with some KBase leads, and Adam raised some good points, but I'd like to bring this discussion project wide to the devel list.
I want to publish a manuscript where we sequenced 124 environmental isolates, resulting in a total of 166 new genomes.
We sequenced the genomes, assembled them in KBase, annotated them, modeled them. We have 124 biolog arrays for all sample. I want KBase to be my supplementary material. I want to send people to KBase to get all this data and analysis, rather than pointlessly dumping them into NCBI where they won't be connected to any of the analysis we did (plus I don't really want to jump through the hoops that NCBI would probably make me jump through).
So how do we support this?
To get things started, here's a few things I think we need:
1.) Persistent KBase-wide IDs for the genomes and phenotype sets that will never change I'm suggesting "kb|g.###" for genomes, and "kb|g.###.phenos.###" for phenotype sets. Comments?
2.) A place people can go to look at the data, download data, access data Initially I'm suggesting I put the genomes and all derived data in one or more workspaces, and put links to the workspace in the workspace browser in the publication, then we make sure the workspace offers good mechanisms for viewing and searching all these phenotypes. Comments?
If we are going to link in to workspaces then they need to be made immutable .. Which means there needs to be a transition from a user workspace to an archived reference workspace/data store. Once that transition happens its a reference bel object and "owned" by kbase not the user. It becomes a public dataset etc.
Genomes if they are public should be in the CDM and treated the same as other public genomes.
(Btw I'm discussing with jgi the issue of publishing genomes without going through NCBI. They wish to do this as well since they can no longer afford the overhead of iterating with NCBI on genome submission) ..
Jim Bristow and I would like to do this in a coordinated way starting ona particular date.
3.) We need to decide on what metadata we want to support for the release of genomes, models, and Biolog array data in KBase I don't have a strong proposal here. I'm hoping some of the metadata experts speak up.
You need something like the MIGS, MIM and MIBA (I made up the last two) standards. There are two aspects of this. The information we require for these datasets and the metadata mechanism.
I strongly believe we need a single metadata mechanism for all of kbase that works for data in all of the stores (rather than one for each store or even worse a different mechanism for each type).
Folker has proposed a mechanism that with some thought could be adopted as the scheme for all the stores. Each core data type needs to have a well defined set of items they want for each data set. I believe the mechanism proposed can support that.
4.) What are the things about KBase that we find to be the most embarrassing… that we would absolutely want to correct before any manuscript like the one I'm proposing gets published and starts driving people enmasse to KBase. I'm thinking documentation at the moment… and I'm the number one offender here. We have shining spots in the documentation, but we've got some real sore spots too.
Documentation for sure... Needs to be updated and made consistent with the current releases of tools etc.
This is something I think we want to address in short order. If we shift to campaigns, I think we should have a campaign specifically to address these issues. Because if a user of KBase cannot easily publish the work they do in KBase right now, then we're going to have trouble attracting users. The best possible advertisement is a good manuscript that people read and get excited about all the powerful things they can do in KBase. Then they click on our supplementary material and come to KBase and see the tools in action.
Chris _______________________________________________ Kbase-devel mailing list [email protected] https://lists.kbase.us/mailman/listinfo/kbase-devel