Skip to content

Inconsistency in protocol buffers specification? #28

Description

@Gabriella439

The protocol buffers encoding specification says:

Normally, an encoded message would never have more than one instance of a non-repeated field. However, parsers are expected to handle the case in which they do. For numeric types and strings, if the same field appears multiple times, the parser accepts the last value it sees. For embedded message fields, the parser merges multiple instances of the same field, as if with the Message::MergeFrom method – that is, all singular scalar fields in the latter instance replace those in the former, singular embedded messages are merged, and repeated fields are concatenated. The effect of these rules is that parsing the concatenation of two encoded messages produces exactly the same result as if you had parsed the two messages separately and merged the resulting objects. That is, this:

MyMessage message;
message.ParseFromString(str1 + str2);

is equivalent to this:

MyMessage message, message2;
message.ParseFromString(str1);
message2.ParseFromString(str2);
message.MergeFrom(message2);

Let's call this a monoid homomorphism law for ParseFromString

However, this seems to conflict with another protocol buffers design decision to omit default fields when serializing message. Specifically, the protocol buffers 3 language guide says:

When a message is parsed, if the encoded message does not contain a particular singular element, the corresponding field in the parsed object is set to the default value for that field.

Defaulting values breaks the above monoid homomorphism law if you consider the following message type:

message Example {
  int32 foo = 1;
}

... and then you serialize the following two messages:

  • message1 : an Example message with the field foo set to 1
  • message2: an Example message with the field foo set to 0

If you deserialize each ByteString independently and combine the resulting Messages you would get the field foo set to 0 (since the latter Message would default to foo = 0 and fields from the latter Message take precedence), but if you combined the ByteStrings before deserializing, you would get the field foo set to 1 (since the 0 field was never explicitly serialized)

Is there something that I'm missing or is the specification inconsistent?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions