Replies: 2 comments 1 reply
|
What if we added each author to each triple to figure out who worked on a set of edits? from @shaps80 |
0 replies
|
Time-ordered UUIDs for CRDT-y things? from @shaps80. If you create a lot of ops at once programatically only using a timestamp (in ms) on the op/triple, the timestamp may not be small enough to distinguish the order since many ops might have the same timestamp. |
1 reply
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
This post describes the new data model for offchain data being introduced in the first public version of the Geo knowledge graph. This doc covers the data that gets stored on our decentralized storage (IPFS), and not the data queryable from the data service.
Motivation
The goal of Geo's data model is to create a graph-like structure that can represent arbitrary relationships between "things." There's traditionally been many approaches for doing this: knowledge graphs, labeled property graphs, RDF stores, and more. Geo combines elements from many of these approaches into a simple data structure called a "Triple." Each Triple in Geo composes together to form the Geo knowledge graph.
v1.0.0 of the Geo knowledge graph formalizes the spec and defines how these Triples should be posted onto IPFS. The data on IPFS then gets turned into queryable data by the Geo data service for applications building on top of the knowledge graph. You can read about the knowledge graph and data service architectures here.
Triples and Ops
A Triple is represented on IPFS as an "Op" (short for Operation) which describes the change being made to a Triple. In v1.0.0 of the spec, a Triple can either be "SET" or "DELETE(d)." A Triple is unique in the knowledge graph based on the combination of its Space, Entity, and Attribute identifiers, i.e., only one triple with a given (Space, Entity, Attribute) can exist. This means that all changes to a Triple are automatically cast as "upserts" in the data service.
The "Op" data structure looks like this:
Native value types
A Native Value Type in the Geo knowledge graph refers to the "type" of a value. These are Entity, Text, Number, URL, Checkbox, and Collection. Each NVT provides semantic meaning as to what the value is. An application might provide unique UX for each of these NVTs and the data service might provide additional APIs based on the value type like sorting, full-text search, and more. Each NVT value is stored directly on a Triple's value. Read about the knowledge graph architecture here.
Rendering different entity types
Additionally, users are able to define the "type" of an entity within the knowledge graph itself. The type of an entity might signal to render this entity in a specific way for specific applications.
For example, Images in the knowledge graph are they themselves entities. So to represent an Image as a value in a triple you are actually referencing another entity. To render this entity as an image we need to know that it's an image at render-time.
v1.0.0 defines a set of "Universally Renderable Entity Types." These are exposed by the data service so applications know that a Triple with an Entity NVT is a specific type of entity that they can render a specific way if they choose. For v1.0.0, we expose "Image" as the first URET. Read about Universal Renderable Entity Types here.
[WIP] Value options
We had discussed an optional "options" field on an Op's value. This design needs to be fleshed out.
Collections
Since a given Triple's Space, Entity, Attribute combination is now unique in the knowledge graph, it's no longer possible to associate many values with a given (S, E, A). Collections are a new data model in Geo meant to solve this. To associate many values with a given (S, E, A) you create a Triple with a Native Value Type of "COLLECTION."
Collections are a powerful new mechanism for creating lists of entities in Geo. Read more about Collections here.
Writing to the knowledge graph and IPFS action types
v1.0.0 of the knowledge graph spec defines several data structures representing the different "action types" that get posted onto IPFS for use in the system. These includes a set of edits, requesting space permissions, creating proposals, and more. For the sake of this doc we'll cover the Edit action type since it's the most relevant for the offchain data spec. Read this doc to learn about IPFS action types.
v1.0.0 defines an Edit as a set of Ops applied to a space. An Edit also contains additional metadata like the authors of the Edit, a unique identifer, a name, and the version of the action type for backwards compatibility.
Important
The Geo knowledge graph expects that this Op is encoded into a binary format using Protocol Buffers before being posted onto IPFS. The data service decodes the Protobuf definition back into the expected data structure at index time. The goal of using Protobufs is to drastically reduce the size of the data posted onto decentralized storage.
The Geo SDKs provide APIs for encoding the different IPFS action types before posting on IPFS. Read this doc for more information on the different IPFS action types and our binary encodings for each in the Geo knowledge graph.
Identifiers
The Geo knowledge graph is effectively a giant, decentralized graph database, so data in the knowledge graph is highly referential. Triples reference attributes by their identifier, and entity values reference the identifier of the entity being referenced.
v1.0.0 of the knowledge graph defines all identifiers in Geo as Universally Unique Identifiers (UUID) Version 4 with the hyphens (-) removed. The goal of removing hyphens is to reduce the amount of data stored on IPFS by 4 bytes for every identifer. Since Geo is highly referencial, this should add up quite a bit over time.
Breaking changelog
All reactions