github dolthub/dolt v2.4.0
2.4.0

3 hours ago

Merged PRs

dolt

  • 11956: Fix upsert authorization and add protocol regression coverage
    Upserts could rewrite rows without UPDATE privileges. Update GMS to require UPDATE authorization and add protocol regression coverage for denied and permitted upserts.
  • 11935: Replace github.com/dolthub/fslock with github.com/dolthub/file-locks
    The new library keeps the interface - New, Lock, TryLock, LockWithTimeout, Unlock, Close, ErrLocked and ErrTimeout - so nothing here changes but the import line. The package is still named fslock, imported under that name from a path that no longer matches it.
  • 11933: Allow DoltgreSQL to override effective user in search path $user substitution.
  • 11921: Allow DoltgreSQL to provide branch-control principal
  • 11920: Add DoltgreSQL transaction lifecycle hooks
    These allow Doltgres to see transaction start, why a transaction ended (rollback, commit) and interact with savepoints as well.
  • 11896: Limit CI deferral comments to label and draft state changes
    Limits CI deferral and release comments to relevant label changes and draft/ready transitions, preventing duplicate notifications on commit pushes.
  • 11893: go/store/nbs: Fix two nil-ghostGen bugs in GenerationalNBS
    Two pre-existing bugs on the ghostGen == nil path, both only reachable outside the Dolt server (dbfactory always supplies a ghost store):
    • HasMany returned nil instead of the absent set, hiding missing hashes. Regression test included.
    • PersistGhostHashes had its nil check inverted.
  • 11892: go/store/nbs: Fix generation ordering on reads to avoid missing a chunk which is being moved from new to old gen.
    When GC moves chunks from new to old gen, it publishes them into old gen and then, eventually, drops them from new gen. If the Reads check old gen first, and then fall back to new gen, they can miss a chunk which is actually retained --- the entire GC can run in between the return from the old gen Read and before the new gen Read.
    Reading new gen first is the fix. If the chunk isn't in new gen and it does exist, it has guaranteed already been published to old gen by the time we ask for it.
  • 11891: go/store/nbs: Drain outstanding reads before a GC installs its keeper.
    An unlocked read samples |nbs.keeperFunc| under |nbs.mu| and then runs
    against a table set with the lock released, so a read already in flight
    when BeginGC installed its keeper could return chunks having taken no
    read dependency on them, leaving the GC free to collect them. BeginGC now
    sets |gcInstallPending| and waits for |outstandingReads| to drain before
    installing, and every read path parks on that flag at the top of its
    locked section. Parking there, rather than in |beginRead|, keeps a read
    which spans the memtable and the table files from being split across the
    barrier. A BeginGC whose context is cancelled during the drain clears the
    flag and restores conjoin, so it leaves no reads parked and no conjoin
    disabled.
  • 11889: go/store/nbs: Sample the keeper in |beginRead|, and register every unlocked read.
    |beginRead| now samples |nbs.keeperFunc| and |nbs.gcCycleCounter| itself
    and registers every unlocked read, rather than each call site sampling
    them separately and only reads which started during a GC being
    registered. A read's keeper and cycle therefore come from the same
    locked section which registers it, |nbs.outstandingReads| counts every
    read running against a table set, and |endRead| is never nil. This is
    the plumbing for making a keeper installation a barrier: the follow-up
    drains those registrations before it installs a keeper. The costs are
    that a read with no GC running reacquires |nbs.mu| once at completion,
    where before it did not, and that EndGC now also waits out reads which
    started before its GC began.
  • 11888: go/store/nbs: Fix a lost cancellation wakeup in |waitForGC|.
    Now take nbs.mu when we broadcast gcCond, so that our waiter is definitely already parked if they need to see our wakeup.
  • 11887: go: NodeStore, ValueStore: Avoid caching things that might have been read before an incoming Purge request.
    The NodeStore and the ValueStore cache have the structure where they are queried and, if there is a miss, a fetch is made the underlying ChunkStore. They are then populated with the result. GC runs a Purge operation on these caches. This makes all future read operations reach the ChunkStore layer during the GC operation, so that proper dependencies can be taken on them.
    This change fixes a race where Get -> Read -> Purge -> Put could result in the cache carrying data throughout and after the GC which might have been read before the Purge operation itself. The caches now carry a purge count and they refuse to populate with values which were fetched based on a Get at a previous purge count.
  • 11872: integration-tests/go-sql-server-driver: Add machinery to run dolt fsck on the database directories anytime these tests fail.
  • 11854: Print SQL boolean results as numeric values
    Prints boolean SQL results as 1 and 0 in text and CSV output to match MySQL.
    Fixes #6044
  • 11852: Validate destination branch names before pushing
    Rejects invalid destination branch names before pushing to remotes.
    Fixes #5341
  • 11851: Enable Dolt executable comments in SQL scripts
    Enables DOLT executable comments with the parser support in dolthub/vitess#494.
    Fixes #3096
  • 11849: Support JSON Lines output for SQL queries
    Adds jsonl SQL result formatting with one JSON object per row.
    Fixes #1844
  • 11844: Show ON UPDATE CURRENT_TIMESTAMP across schemas and table alterations in EXTRA column
    Fix #11774
    Blocked by dolthub/go-mysql-server#3862
  • 11773: go/store/nbs: Add validation and sanity checks on the index data when a NomsBlockStore opens a table file.
    Two types of checks are added. Simple checks are run on every index parse. They assert that prefixes are in sorted order, that ordinals are within range, that offsets (dervied from the lengths) are not too small. The more expensive checks only run when we add new table files to the store, typically as part of GC, a pull or receiving a push. These further assert that recorded ordinals are unique and that the hash of the suffixes table matches the file name, which is how the table file name is derived on all write paths.
    These sanity checks are added to make Dolt more robust in the face of bugs in Dolt or unexpected behavior from I/O interactions. If these validation checks fail, Dolt has a chance to fail a write operation which is attempting to take a new dependency on the corrupted data.
  • 11643: dynamic merge conflict handling
    This change allows for more flexible row merge behavior, in particular the ability to detect conflict on matching values and create rows dynamically based on the callers requirements.
    DumboDB is the direct consumer of this change. This allows us to match Dumbo's standard compare-and-set query pattern, and to also enable collection specific merge modes.
    There may be cleaner ways to implement this, but I opted to touch the common (dolt) code path as little as possible since this could have significant perf impacts. nil rowMergePolicy is the dolt way as a result.
  • 11618: {proto,go/serial}/.bazelversion: Bump to 9.2.0. Bump some module versions.

go-mysql-server

  • 3946: Default to InternalDecimalType for CONVERT with decimal arguments
    Fix regression from dolthub/go-mysql-server#3939
    dolt bump in #11915
    doltgres bump in dolthub/doltgresql#3421
  • 3940: Require UPDATE privileges for duplicate-key inserts
    INSERT ... ON DUPLICATE KEY UPDATE could rewrite rows without UPDATE privileges. Require UPDATE authorization and return the MySQL-compatible error code when access is denied.
  • 3939: fix panic for CAST and CONVERT to types with invalid precision/scale
    Ideally, CAST/CONVERT should only require the destination sql.Type to make the proper conversion.
    However, a significant portion of our conversion and comparison logic is so dependent on existing hackiness that it would require an entire rewrite to properly implement.
    As a result, this is a minimal change that fixes the panic and leaves the logic in a slightly better place.
  • 3924: Fix panic on nil decimal values in VALUES derived tables
    Fix #11942
  • 3923: Resolve untyped NULL columns in CTAS to VARBINARY(0)
    Fix #11941
  • 3921: Fix non-literal evaluation in unix_timestamp
    Fix #11917
  • 3920: Fix panic converting strings with a dangling exponent
    Fix #11918
  • 3919: implement precision for TIME types
    Add support for precision parameter for TIME types.
    Implements proper rounding and string output according to the precision.
    fixes: #10661
  • 3911: Add delimited time regex
    Fixes a regression from #3901 that breaks the dolt bats test import-create-tables: table import -c infers types from data (#11922).
  • 3908: Unwrap ANY_VALUE for aggregates and window functions
    Fix #11912
  • 3907: Optimize aggregate functions and remove sort nodes on multi-column indexes provided that any columns before the aggregated/sorted column are constant
    Dolt attempts to remove ORDER BY nodes from the plan when the chosen index is guaranteed to already produce rows in the desired order. But currently, it's limited to cases where the sort expressions are a prefix of the index. This means that we don't currently remove the sort in examples like this one:
    CREATE TABLE test (pk1 int, pk2 int);
    SELECT pk2 FROM test WHERE pk1 = 64 ORDER BY pk2 LIMIT 1;
    
    We can optimize this by using the primary index, because even thought we're not ordering by pk1, the value of pk1 is constant for all results.
    Additionally, Dolt tries to optimize MIN and MAX functions into the above. This also currently requires that that the column being aggregated is the first column in the index, meaning that we can't optimize the following:
    CREATE TABLE test (pk1 int, pk2 int);
    SELECT MIN(pk2) FROM test WHERE pk1 = 64;
    SELECT MIN(pk2) FROM test WHERE pk1 = 64 GROUP BY pk1;
    
    The second SELECT is unlikely to be written by hand but could be produced procedurally or as the result of other optimizations.
  • 3904: Fix ONLY_FULL_GROUP_BY validation for window functions
  • 3901: handle no delimiter datetime string parsing
    Fix support for parsing non-delimited datetime strings to DATE, DATETIME, and TIMESTAMPS.
    Additionally, fixes parsing for weirdly delimited datetime strings.
    Fixes: #10278
  • 3900: Fixed duplicate window columns
    Fix required for:
  • 3899: Fix dropped residual filters and false branch matching in concat lookup join
    Fix #11886
  • 3898: Use common extended type for VALUES rows
    Delegate VALUES common-type selection to the extended-type hook when both row expressions use extended types. This lets dialect integrations resolve unknown NULL literals against typed VALUES rows while preserving existing built-in type handling.
    Includes focused schema inference coverage.
    Part of dolthub/doltgresql#3388
  • 3897: Optimize SPACE with strings.Repeat
    Fix #11408
  • 3889: Filled generation_expression and extra for generated columns in information_schema.columns
    information_schema.columns now reports the generation expression and STORED/VIRTUAL GENERATED extra for generated columns instead of leaving them empty. Although this was found and fixed in the Doltgres integrator, GMS requires a different fix as Doltgres uses its own information_schema table implementation.
  • 3885: Support wildcard replication filter values
    Adds native string-list values for REPLICATE_WILD_DO_TABLE and REPLICATE_WILD_IGNORE_TABLE, including MySQL-compatible format validation and SHOW REPLICA STATUS formatting.
    Simplifies replication option handling to use native Go values instead of exported wrapper types.
    Companion Dolt PR: github.com//pull/11842
    Companion: dolthub/dolt#11842
  • 3867: Resolve order by columns using column ID instead of name to avoid ambiguity
    fixes #11418
  • 3862: Always generate format column ON UPDATE clauses in EXTRA
    Fix EXTRA column to always generate on DESCRIBE, SHOW COLUMNS, and information_schema.columns.
    Block #11844

vitess

  • 494: Support Dolt executable SQL comments
    Parses DOLT executable comments for the SQL integration in #11851.
    Fixes #3096
  • 493: Add wildcard replication filter syntax
    Adds parser and AST support for REPLICATE_WILD_DO_TABLE and REPLICATE_WILD_IGNORE_TABLE in CHANGE REPLICATION FILTER statements. Wildcard values are represented as ordered quoted-string lists, including support for explicitly empty lists.
    Parser coverage includes multiple patterns, wildcard escapes, quotes, Unicode, empty lists, and rejection of unquoted patterns.
    Part of #11787

Closed Issues

  • 11418: Dolt rejects a repeated window expression in ORDER BY as ambiguous.
  • 11942: Dolt panics on a NULL DECIMAL in a VALUES-derived table
  • 11941: Dolt panics when CTAS materializes an untyped NULL
  • 11917: Server panic kills the connection on SELECT UNIX_TIMESTAMP(IFNULL(SIN(WEEKDAY(UUID())), (SELECT 1)))
  • 10661: time precision is not respected
  • 11918: Server panic in decimal conversion on aggregate query with WHERE ROUND(HEX(int_col))
  • 11912: ANY_VALUE nested inside another aggregate fails: misleading table not found: <alias> (1146) and internal unable to find field with index N in row of M columns. This is a bug. (1105)
  • 11886: Lookup join drops an AND conjunct when the ON clause also has an OR over indexed columns (wrong results)
  • 11408: SPACE(n) results in a hang when n is large
  • 11490: Wrong results / internal error: a window function + correlated scalar subquery corrupts CREATE TABLE AS (persists 0/NULL) and crashes INSERT .. SELECT
  • 11774: bug/compatibility-mysql: "information_schema.columns.EXTRA" omits "on update CURRENT_TIMESTAMP"
  • 6044: dolt returns boolean values sometimes
  • 6612: DOLT_COMMIT_DIFF_table is missing data when there's a schema collation change merged in
  • 3096: Dolt should support it's own flavor of special comments.
  • 7905: JSON functions return incorrect values when an object key path is used on an array value.
  • 6236: dolt_checkout() to change branches doesn't work in stored procedures
  • 6318: Checks when creating/altering a FK need to correctly handle revision DBs
  • 8323: Incorrect results for virtual column in WHERE or ORDER BY
  • 5341: push should validate branch name like checkout or branch
  • 1844: Allow jsonl result format
  • 8389: Exporting line breaks interrupts CSV formatting
  • 8388: Exporting table replaces empty string with NULL values
  • 6500: Missing insert source aliases parsing
  • 6152: dolt_procedures table does not exist until a procedure is created
  • 4501: GRANT wildcard support
  • 6918: the row result is not printed when there is IF clause used in the stored procedure when using sql server
  • 7196: JSON ordering is not well-ordered.
  • 6742: DECLARE CONTINUE HANDLER failure in dolt
  • 4233: order of CTEs defined matters

Don't miss a new dolt release

NewReleases is sending notifications on new releases.