Feature Description (功能描述)
Add an arbitrary-precision decimal property type: DataType.DECIMAL backed by java.math.BigDecimal, with REST data_type: DECIMAL and PropertyKey.Builder.asDecimal().
Motivation
The numeric property types today are BYTE/INT/LONG/FLOAT/DOUBLE. Values that do not fit a long and must not be rounded (token balances in wei go up to 2^256 - 1, 78 digits; money amounts in general) can only be stored as TEXT. That loses the one place where the server itself does arithmetic: update_strategies in PUT /graph/{vertices,edges}/batch (SUM / BIGGER / SMALLER). UpdateStrategy already computes in BigDecimal, but the result goes back to the property's type, so with DOUBLE a SUM of 10^18 + 1 is 10^18; TEXT fails the strategy's Number type check. For accumulating balances during an import this blocks the use case.
Proposal (points to confirm)
- Type code
12 in DataType (server and the hugegraph-struct copy), name "decimal". Please confirm the code is free.
- No sort key, no index, no OLAP range.
isNumber() stays false; the schema builders reject these with an explicit message. Reason: there is no fixed-width byte-order-preserving encoding for a decimal, and faking one through LongEncoding would be lossy. SUM/MAX/MIN aggregate types on the property key are allowed as for numbers.
- Encoding in
BytesBuffer: vint(len) + unscaled two's-complement bytes + vint(scale). Exact for any precision, 33 bytes for a uint256, scale preserved; existing encodings untouched.
- JSON: always a string on output (
toPlainString()), a string or a number literal accepted on input. A JSON number is a double to most clients; a string is the only lossless representation. In the batch update a fraction has to be sent as a string, because the request's properties map is parsed by Jackson before any schema is known (a 0.000000000000000001 literal becomes a double); integral literals are exact.
Also in the same change: exact comparison in ConditionQuery when one side is a BigDecimal (instead of through doubleValue()), a string variant in the store-side row decoder, and one fix in BatchAPI.updateExistElement: the JSON value is normalised through the property key before the strategy runs, on both paths. Today the strategy receives the raw JSON value, which only works for the types Jackson happens to produce; a decimal (or a date) sent as a string fails the type check.
Status
Implementation and tests are ready on the branch feat/decimal-datatype in my fork (one commit on top of master 60c8803, 25 files): unit tests (DataTypeTest, BytesBufferTest, JsonUtilTest), core (PropertyKey/IndexLabel/EdgeLabel/VertexCoreTest), API (VertexApiTest: SUM on 2^256-2 + 1, two entries of one vertex within one request, BIGGER) and struct. Run locally with the CI scripts: unit, core on rocksdb and memory, api on rocksdb, all green. I will open the PR tomorrow to leave time for comments on the points above; the client/loader in hugegraph-toolchain will be a separate change.
If any of the four points looks wrong, I would rather hear it here than in the code review.
Feature Description (功能描述)
Add an arbitrary-precision decimal property type:
DataType.DECIMALbacked byjava.math.BigDecimal, with RESTdata_type: DECIMALandPropertyKey.Builder.asDecimal().Motivation
The numeric property types today are
BYTE/INT/LONG/FLOAT/DOUBLE. Values that do not fit alongand must not be rounded (token balances in wei go up to 2^256 - 1, 78 digits; money amounts in general) can only be stored asTEXT. That loses the one place where the server itself does arithmetic:update_strategiesinPUT /graph/{vertices,edges}/batch(SUM/BIGGER/SMALLER).UpdateStrategyalready computes inBigDecimal, but the result goes back to the property's type, so withDOUBLEaSUMof10^18 + 1is10^18;TEXTfails the strategy'sNumbertype check. For accumulating balances during an import this blocks the use case.Proposal (points to confirm)
12inDataType(server and thehugegraph-structcopy), name"decimal". Please confirm the code is free.isNumber()staysfalse; the schema builders reject these with an explicit message. Reason: there is no fixed-width byte-order-preserving encoding for a decimal, and faking one throughLongEncodingwould be lossy.SUM/MAX/MINaggregate types on the property key are allowed as for numbers.BytesBuffer:vint(len)+ unscaled two's-complement bytes +vint(scale). Exact for any precision, 33 bytes for a uint256, scale preserved; existing encodings untouched.toPlainString()), a string or a number literal accepted on input. A JSON number is adoubleto most clients; a string is the only lossless representation. In the batch update a fraction has to be sent as a string, because the request'spropertiesmap is parsed by Jackson before any schema is known (a0.000000000000000001literal becomes adouble); integral literals are exact.Also in the same change: exact comparison in
ConditionQuerywhen one side is aBigDecimal(instead of throughdoubleValue()), a string variant in the store-side row decoder, and one fix inBatchAPI.updateExistElement: the JSON value is normalised through the property key before the strategy runs, on both paths. Today the strategy receives the raw JSON value, which only works for the types Jackson happens to produce; a decimal (or a date) sent as a string fails the type check.Status
Implementation and tests are ready on the branch
feat/decimal-datatypein my fork (one commit on top ofmaster60c8803, 25 files): unit tests (DataTypeTest,BytesBufferTest,JsonUtilTest), core (PropertyKey/IndexLabel/EdgeLabel/VertexCoreTest), API (VertexApiTest:SUMon 2^256-2 + 1, two entries of one vertex within one request,BIGGER) and struct. Run locally with the CI scripts: unit, core on rocksdb and memory, api on rocksdb, all green. I will open the PR tomorrow to leave time for comments on the points above; the client/loader in hugegraph-toolchain will be a separate change.If any of the four points looks wrong, I would rather hear it here than in the code review.