Skip to content

[fix](statistics) Treat invalid column statistics as UNKNOWN instead of aborting analyze job - #66756

Open
codeDing18 wants to merge 2 commits into
apache:masterfrom
codeDing18:statistic-null
Open

[fix](statistics) Treat invalid column statistics as UNKNOWN instead of aborting analyze job#66756
codeDing18 wants to merge 2 commits into
apache:masterfrom
codeDing18:statistic-null

Conversation

@codeDing18

Copy link
Copy Markdown
Contributor

What problem does this PR solve?

Issue Number: close #64122

Related PR: #xxx

Problem Summary:

Release note

None

Check List (For Author)

  • Test

    • Regression test
    • Unit Test
    • Manual test (add detailed scripts or steps below)
    • No need to test or manual test. Explain why:
      • This is a refactor/code format and no logic has been changed.
      • Previous test can cover this change.
      • No code files have been changed.
      • Other reason
  • Behavior changed:

    • No.
    • Yes.
  • Does this need documentation?

    • No.
    • Yes.

Check List (For Reviewer who merge this PR)

  • Confirm the release note
  • Confirm test cases
  • Confirm document
  • Add branch pick label

@hello-stephen

Copy link
Copy Markdown
Contributor

Thank you for your contribution to Apache Doris.
Don't know what should be done next? See How to process your PR.

Please clearly describe your PR:

  1. What problem was fixed (it's best to include specific error reporting information). How it was fixed.
  2. Which behaviors were modified. What was the previous behavior, what is it now, why was it modified, and what possible impacts might there be.
  3. What features were added. Why was this function added?
  4. Which code was refactored and why was this part of the code refactored?
  5. Which functions were optimized and what is the difference before and after the optimization?

@codeDing18

Copy link
Copy Markdown
Contributor Author

run buildall

englefly
englefly previously approved these changes Aug 14, 2026
@github-actions github-actions Bot added the approved Indicates a PR has been approved by one committer. label Aug 14, 2026
@github-actions

Copy link
Copy Markdown
Contributor

PR approved by at least one committer and no changes requested.

@github-actions

Copy link
Copy Markdown
Contributor

PR approved by anyone and no changes requested.

@englefly

Copy link
Copy Markdown
Contributor

建议改一下标题
fix Treat invalid column statistics as UNKNOWN instead of aborting analyze job

@hello-stephen

Copy link
Copy Markdown
Contributor

FE Regression Coverage Report

Increment line coverage 1.52% (7/462) 🎉
Increment coverage report
Complete coverage report

@codeDing18

Copy link
Copy Markdown
Contributor Author

建议改一下标题 fix Treat invalid column statistics as UNKNOWN instead of aborting analyze job

OK

@codeDing18 codeDing18 changed the title [fix](statistics) fix ColStatsData.isValid() falsely rejects sampled column statistics when a column is (almost) all NULL [fix](statistics) Treat invalid column statistics as UNKNOWN instead of aborting analyze job Aug 14, 2026
@codeDing18

Copy link
Copy Markdown
Contributor Author

run buildall

@github-actions github-actions Bot removed the approved Indicates a PR has been approved by one committer. label Aug 14, 2026
@hello-stephen

Copy link
Copy Markdown
Contributor
TPC-H: Total hot run time: 17700 ms
machine: 'aliyun_ecs.c7a.8xlarge_32C64G'
scripts: https://github.com/apache/doris/tree/master/tools/tpch-tools
Tpch sf100 test result on commit fc5d3a15d474c94c9ecca3596c37894a09d8be7c, data reload: false

------ Round 1 ----------------------------------
============================================
q1	17572	3032	3020	3020
q2	1939	231	156	156
q3	10409	887	520	520
q4	4668	247	201	201
q5	7673	577	390	390
q6	134	115	101	101
q7	550	508	395	395
q8	9243	956	935	935
q9	3459	2373	2385	2373
q10	6535	900	718	718
q11	438	247	235	235
q12	682	400	327	327
q13	17876	1863	1539	1539
q14	159	151	141	141
q15	q16	446	408	375	375
q17	760	795	873	795
q18	3167	2324	2294	2294
q19	1103	852	820	820
q20	620	530	461	461
q21	5305	1665	1944	1665
q22	327	275	239	239
Total cold run time: 93065 ms
Total hot run time: 17700 ms

----- Round 2, with runtime_filter_mode=off -----
============================================
q1	3428	3360	3334	3334
q2	210	207	156	156
q3	2259	2292	2244	2244
q4	1200	1176	894	894
q5	2199	2157	2171	2157
q6	178	121	89	89
q7	1071	940	883	883
q8	1602	1392	1396	1392
q9	3088	3081	3076	3076
q10	1881	1841	1663	1663
q11	348	274	251	251
q12	468	438	345	345
q13	1806	1879	1553	1553
q14	183	182	168	168
q15	q16	402	395	360	360
q17	1032	1039	1031	1031
q18	5018	4492	4795	4492
q19	857	856	877	856
q20	959	946	826	826
q21	3897	3155	3265	3155
q22	451	405	345	345
Total cold run time: 32537 ms
Total hot run time: 29270 ms

@hello-stephen

Copy link
Copy Markdown
Contributor
TPC-DS: Total hot run time: 86534 ms
machine: 'aliyun_ecs.c7a.8xlarge_32C64G'
scripts: https://github.com/apache/doris/tree/master/tools/tpcds-tools
TPC-DS sf100 test result on commit fc5d3a15d474c94c9ecca3596c37894a09d8be7c, data reload: false

query5	4236	411	334	334
query6	400	162	156	156
query7	4889	461	272	272
query8	294	120	114	114
query9	8676	2944	2921	2921
query10	397	253	227	227
query11	5358	1043	942	942
query12	120	74	71	71
query13	1207	453	339	339
query14	5975	2032	1892	1892
query14_1	1806	1776	1785	1776
query15	176	132	113	113
query16	941	401	387	387
query17	805	486	396	396
query18	2359	345	258	258
query19	167	155	117	117
query20	78	72	80	72
query21	210	120	103	103
query22	5577	5448	5381	5381
query23	7318	6858	6572	6572
query23_1	6656	6713	6757	6713
query24	7286	1105	777	777
query24_1	827	791	781	781
query25	407	289	247	247
query26	1243	272	166	166
query27	2707	450	284	284
query28	4638	1531	1512	1512
query29	921	440	345	345
query30	281	182	150	150
query31	966	667	610	610
query32	105	50	49	49
query33	471	207	183	183
query34	986	869	493	493
query35	403	399	328	328
query36	552	535	530	530
query37	125	81	69	69
query38	997	851	843	843
query39	533	538	538	538
query39_1	511	536	532	532
query40	221	154	118	118
query41	54	51	52	51
query42	78	74	75	74
query43	247	252	220	220
query44	1023	571	583	571
query45	114	99	99	99
query46	781	845	542	542
query47	961	1012	947	947
query48	328	307	238	238
query49	531	249	194	194
query50	837	323	265	265
query51	8112	8181	8147	8147
query52	69	69	61	61
query53	209	223	169	169
query54	228	190	173	173
query55	80	58	58	58
query56	244	227	224	224
query57	662	632	617	617
query58	242	200	190	190
query59	1081	1105	1019	1019
query60	304	212	211	211
query61	148	136	134	134
query62	353	202	193	193
query63	182	162	160	160
query64	2716	747	623	623
query65	1576	1589	1605	1589
query66	1804	284	237	237
query67	10108	9725	9922	9725
query68	3067	1184	756	756
query69	347	227	194	194
query70	640	588	609	588
query71	313	257	240	240
query72	2333	1774	1562	1562
query73	662	573	320	320
query74	2000	1233	1140	1140
query75	1245	1172	1013	1013
query76	2391	748	575	575
query77	240	259	205	205
query78	5193	4932	4581	4581
query79	1445	878	589	589
query80	1150	385	344	344
query81	485	197	170	170
query82	638	142	111	111
query83	320	259	232	232
query84	296	122	104	104
query85	844	423	387	387
query86	401	169	169	169
query87	1001	986	907	907
query88	2825	2150	2149	2149
query89	313	238	206	206
query90	1830	150	143	143
query91	159	145	132	132
query92	53	46	47	46
query93	1373	1186	811	811
query94	607	251	241	241
query95	620	426	349	349
query96	874	625	273	273
query97	1118	1101	1040	1040
query98	142	139	131	131
query99	410	352	304	304
Total cold run time: 179669 ms
Total hot run time: 86534 ms

@hello-stephen

Copy link
Copy Markdown
Contributor
ClickBench: Total hot run time: 14.75 s
machine: 'aliyun_ecs.c7a.8xlarge_32C64G'
scripts: https://github.com/apache/doris/tree/master/tools/clickbench-tools
ClickBench test result on commit fc5d3a15d474c94c9ecca3596c37894a09d8be7c, data reload: false

query1	0.00	0.01	0.00
query2	0.08	0.03	0.04
query3	0.25	0.11	0.11
query4	1.60	0.10	0.10
query5	0.16	0.15	0.16
query6	1.25	0.63	0.61
query7	0.03	0.01	0.00
query8	0.05	0.03	0.03
query9	0.29	0.22	0.21
query10	0.35	0.34	0.32
query11	0.15	0.12	0.11
query12	0.16	0.11	0.12
query13	0.32	0.31	0.32
query14	0.47	0.48	0.47
query15	0.37	0.37	0.35
query16	0.21	0.23	0.22
query17	0.74	0.72	0.74
query18	0.19	0.18	0.18
query19	1.26	1.18	1.22
query20	0.01	0.01	0.01
query21	15.43	0.14	0.10
query22	5.09	0.04	0.04
query23	16.17	0.25	0.11
query24	3.00	0.31	0.27
query25	0.12	0.03	0.04
query26	0.78	0.17	0.12
query27	0.03	0.03	0.03
query28	3.69	0.57	0.28
query29	12.42	3.11	2.54
query30	0.26	0.12	0.12
query31	2.76	0.36	0.18
query32	3.52	0.32	0.24
query33	1.37	1.40	1.43
query34	15.39	2.18	1.80
query35	1.78	1.77	1.76
query36	0.46	0.30	0.28
query37	0.06	0.04	0.05
query38	0.04	0.03	0.02
query39	0.03	0.02	0.02
query40	0.11	0.08	0.08
query41	0.08	0.03	0.02
query42	0.03	0.02	0.02
query43	0.03	0.03	0.03
Total cold run time: 90.59 s
Total hot run time: 14.75 s

@hello-stephen

Copy link
Copy Markdown
Contributor

FE Regression Coverage Report

Increment line coverage 6.67% (7/105) 🎉
Increment coverage report
Complete coverage report

@codeDing18

Copy link
Copy Markdown
Contributor Author

/review

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes for four issues: one production cache/persistence ordering defect and three regression-test contract gaps.

Critical checkpoint conclusions

  • Goal and proof: Persisting this extreme sampled row and consuming it as UNKNOWN matches the agreed issue contract, and both conversion entry points were updated. The cache ordering defect means the old value can still be observed, and the changed collection path is not proved deterministically.
  • Scope and focus: The production change is small and focused. The new regression is disproportionately large for the behavior it proves and contains the three test findings below.
  • Concurrency: Analysis-task threads, asynchronous Caffeine loaders, concurrent planners, and follower cache handlers share the statistics lifecycle. No common lock closes the invalidate-before-flush window; this is the blocking production finding. No new lock-order or deadlock issue was found.
  • Lifecycle: I traced sync and registered jobs from collection through cache publication, AnalysisJob buffering/flush, task completion, metadata freshness, reload/preheat, and restart/failover. Aside from the cache publication order, the intended UNKNOWN lifecycle is consistent; no static-initialization or ownership issue applies.
  • Configuration: No product configuration was added. The new test mutates the existing cluster-global enable_auto_analyze value without non-concurrent isolation or restoration.
  • Compatibility: No function symbol, storage format, or FE-BE protocol field changed. Current direct, cache-load, preheat, SHOW, leader, follower, legacy, and Nereids consumers pass through one of the new guards; no supported rolling-upgrade defect was established.
  • Parallel paths: OLAP full/sample analysis, external and plugin-driven callers, direct SET STATS, connector fallback, and the separate partition-statistics representation were checked. No distinct missed production path survived review.
  • Conditional checks: All three ColStatsData.isValid predicates and their boundaries were checked. Returning UNKNOWN at both construction paths is consistent for current consumers; no additional conditional bug was found.
  • Test coverage: The changed unit tests validate conversion only. Random tablet selection plus permissive assertions let the regression pass without exercising the new runQuery no-throw/persist behavior, and there is no latch-controlled test for the cache race.
  • Test results: The changed SHOW, cached SHOW, memo-plan, and invalid_stats.out expectations are consistent with UNKNOWN statistics. Current FE UT, P0 regression, non-concurrent regression, CheckStyle, compile, and other required PR checks are green.
  • Observability: The existing invalid-stat metric and warning identify the accepted UNKNOWN path; no separate logging or metrics defect was found.
  • Persistence: The statistics row and job freshness metadata are persisted through established paths, but cache invalidation happens before the buffered row is durable and is not repeated after flush.
  • Data writes and failure: A successful buffered insert is not atomic with leader/follower cache publication, producing the stale-cache interleaving called out inline. Flush failures otherwise remain failures; no additional transaction or leak issue was found.
  • FE-BE variables: No new variable or execution option crosses the FE-BE boundary.
  • Performance: No distinct production hot-path regression was found. The two-million-row regression and long polling window buy no deterministic producer-path coverage, which is folded into the test-coverage finding.
  • Other issues and user focus: The trailing database drop violates the required drop-before-use lifecycle, and the global-variable mutation leaks across suites. No additional user focus was supplied.

No local builds or tests were run, as required by the review environment; the conclusions above use code-path inspection and the live checks on the reviewed head.

String message = String.format("ColStatsData is invalid, skip analyzing. %s", colStatsData.toSQL(true));
LOG.warn(message);
throw new RuntimeException(message);
// Don't throw: keep writing the row into the statistics table so that the

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Close the cache-reload window before persisting the invalid row — This branch now reaches syncColStats, which invalidates the cache before job.appendBuf has committed the replacement row. A concurrent planner can start the async loader in that gap, read the previous valid row, and install it after the eviction; flushBuffer never invalidates again, so that stale value can survive until the 48-hour refresh. Please publish a blocking UNKNOWN value or move/repeat leader and follower cache publication after the buffered insert succeeds, and cover this interleaving with a latch-controlled test.

sql """drop database if exists regression_test_analyze_sample_almost_all_null"""
sql """create database regression_test_analyze_sample_almost_all_null"""
sql """use regression_test_analyze_sample_almost_all_null"""
sql """set global enable_auto_analyze=false"""

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Isolate and restore the global auto-analyze setting — This ordinary suite sets a cluster-wide variable to false for a long insert/delete/poll run and never restores it. Normal suites run in parallel, so this can disable auto analysis for unrelated tests and leaves the setting changed after completion or failure. Please run this global-state test in the nonConcurrent group and use setGlobalVarTemporary([enable_auto_analyze: false]) { ... } (or otherwise restore the prior value in finally).


def result = sql """show column stats tbl_del_big"""

// k (NOT NULL) always produces valid sampled stats. Whether v also survives

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Exercise the changed collection path deterministically — Tablet selection uses SecureRandom here, and these assertions deliberately accept either a valid or invalid result for v. If the runQuery throw is reintroduced, this suite can still pass whenever the one x row is sampled; the later ALTER ... SET STATS only exercises the separate read/conversion path. Please retain a mocked runQuery test or force a deterministic sample that proves the invalid row does not throw, is flushed, and is then read as UNKNOWN.

contains("v#1 -> unknown(")
}

sql """drop database if exists regression_test_analyze_sample_almost_all_null"""

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P3] Preserve the regression state after the run — The Doris test rule requires cleanup before setup, not after completion, so failures remain debuggable. This suite already drops the database at line 76; please remove this trailing drop so the MoW data and persisted statistics row survive for inspection.

@codeDing18

Copy link
Copy Markdown
Contributor Author

/review

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes for three new, non-duplicate issues: two production correctness/compatibility defects and one FE test-build failure. The four existing live inline threads are treated as duplicate fences and are not repeated here.

Critical checkpoint conclusions

  • Goal and proof: The intended contract is to let analysis finish, persist the anomalous row, and expose it as UNKNOWN. The two new reader guards implement that contract for fully upgraded FEs, but the follower replay race and old-reader behavior keep the end-to-end goal incomplete. The new deterministic unit proof cannot compile in fe-core.
  • Scope and focus: The production edit is small and focused. The unit test introduces an undeclared mocking framework. No additional user-provided focus was supplied.
  • Concurrency: Analysis tasks, Caffeine loaders, follower RPC handlers, and the independent journal replayer participate in this lifecycle. The blocking put addresses the already-reported invalid-row loader race, but the RPC has no replay ordering and turns a transient valid-payload conversion miss into a long-lived UNKNOWN (MAIN-002). No new lock-order or deadlock issue was found.
  • Lifecycle: I traced collection, leader publication, follower publication, buffered persistence, load/refresh, preheat, SHOW, planning, replay, and failover. The affected negative entry lasts until the 48-hour refresh/eviction path; no static-initialization, ownership, or resource-release issue applies.
  • Configuration: No product configuration or dynamic setting was added. The current regression revision isolates the global auto-analyze setting, so the older live thread is not duplicated.
  • Compatibility: The persisted/RPC representation now relies on reader guards absent from the immediately previous FE. During a supported rolling upgrade, failure of the old master can promote an upgraded follower while an old follower remains, exposing numeric invalid statistics through both RPC and shared-table reload (MAIN-003). There is no protocol or metadata-version fence.
  • Parallel paths: Current-version cache load, preheat, direct/cached SHOW, Nereids, OLAP full/sample collection, manual statistics, and external fallback were checked. They consistently suppress intrinsically invalid rows; no separate missed consumer survived review.
  • Conditional checks: The three isValid() predicates match the intended current-reader safety contract. The follower path incorrectly treats intrinsic invalidity and temporary local metadata unavailability as the same UNKNOWN state; data.isValid() provides the needed distinction.
  • Test coverage: Regression expectations cover UNKNOWN consumption, while the producer-path determinism concern remains in an existing live thread. The replacement deterministic unit test imports JMockit without a fe-core dependency or javaagent, so normal FE build/test paths fail before the test can prove the behavior (MAIN-004).
  • Test results: The changed expected results are consistent with UNKNOWN semantics. On the current head, CheckStyle and repository-hygiene checks pass, but no FE compile, FE unit-test, or regression-test result is present. No local build or test was run, as required by this review environment.
  • Observability: The existing invalid-stat metric and warning are adequate for the accepted UNKNOWN path; no additional logging or metric issue was found.
  • Persistence and data writes: The row uses the existing statistics-table buffer/insert path. The remaining defects are the unordered follower cache publication and mixed-version interpretation after persistence/failover; no additional transaction, crash-leak, or lock issue was established.
  • FE-BE propagation: No new FE-BE variable, Thrift option, or execution flag was added.
  • Performance: No material CPU, memory, or algorithmic regression was found in the added validation. A wrongly retained UNKNOWN can still deprive the follower optimizer of valid statistics for the cache refresh interval, as described inline.
  • Other issues and completion: All ten changed files and relevant adjacent paths were reviewed. After three rounds, every final-round reviewer returned NO_NEW_VALUABLE_FINDINGS; the review is converged, with no unresolved candidate beyond these three comments and the existing live threads.

// Publish a blocking UNKNOWN instead of invalidating, mirroring the leader-side
// syncColStats: invalidation would let a concurrent async loader re-install the
// previous stale row before the leader's buffered insert becomes visible here.
Env.getCurrentEnv().getStatisticsCache().updateColStatsCache(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Keep transient follower conversion misses retryable — toColumnStatistic() returns the same UNKNOWN both for the invalid-stat shape and when this follower has not replayed the referenced column/index yet. The stats RPC has no journal barrier, so a valid update can arrive first; this unconditional put then leaves UNKNOWN cached after metadata catches up, with no replay invalidation until the 48-hour refresh. This is distinct from the existing invalid-row race. Please publish a blocking UNKNOWN only when !data.isValid(); keep other conversion failures invalidated/retryable, and test RPC-before-schema-replay recovery.

String message = String.format("ColStatsData is invalid, skip analyzing. %s", colStatsData.toSQL(true));
LOG.warn(message);
throw new RuntimeException(message);
// Don't throw: keep writing the row into the statistics table so that the

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Fence persisted invalid rows from old FE readers — The new contract depends on both new isValid() guards, but this branch now sends and persists the contradictory row while a base-version FE still converts it into ordinary statistics. During rolling FE upgrade, an upgraded follower can become master after failover while an old follower remains; that old FE then caches the RPC payload and later reloads the same bogus row from the shared table. This is distinct from the existing pre-flush cache-race thread on this line. Please gate persistence until all FEs support the reader contract or use a backward-compatible representation/version fence, and cover the immediately previous reader behavior.

import org.apache.doris.thrift.TUniqueId;

import com.google.common.collect.Lists;
import mockit.Expectations;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Use fe-core's declared mocking framework — These are the only real mockit.* imports in fe-core, but this module declares only mockito-core/mockito-inline; neither its POM nor the FE parent provides org.jmockit:jmockit, and Surefire has no JMockit javaagent. The new deterministic test therefore cannot compile/run in the FE test target. This is distinct from the existing request for deterministic coverage: the supplied replacement itself is unbuildable. Please express it with the already-declared Mockito stack (the removed test already used MockedConstruction), or explicitly add and configure JMockit for this module.

@codeDing18

Copy link
Copy Markdown
Contributor Author

run buildall

@hello-stephen

Copy link
Copy Markdown
Contributor

FE UT Coverage Report

Increment line coverage 37.50% (3/8) 🎉
Increment coverage report
Complete coverage report

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] ColStatsData.isValid() falsely rejects sampled column statistics when a column is (almost) all NULL

4 participants