#### Description
Phase 2 of #6143: semantic-convention families resolve on logs and
metrics, behind the `resolve_semconv_families` flag (default off), on
the storage contract of #12802.
- Registry: each member carries the scope of its own rename edges, so a
fan-out keeps one membership per target and an ambiguous name stays
literal.
- Logs and metrics families need no family code of their own. The gate
applies to every signal, and `LogicalRead` merges members through each
storage's `Read`.
- Metric-name families union the storage names in every `metric_name`
filter, and the querier reads type, temporality, and the reduced flag
across the family.
- Span-metrics labels: the metrics the processor emits, listed by name,
also read each family member with the `resource_` prefix. Requested
names are never rewritten.
- Values suggestions and related values cover every spelling of the
family.
- `deployment.environment.name` resolves on all three signals.
`db.system.name` stays off until a value-mapping reader exists.
#### Additional Information
- `pkg/semconv.Family` fields are now unexported, and `transition.go` is
removed. #12446 reads the old API and needs an update when stacked.
- A target that emits both names of a metric-name family double-counts
in `sum()` during the overlap window. Reading both names is the feature.
Pinned by a test.
- A family of metrics labels keeps the keyless contract of a single
label: no guard, no NULL group.
#### Description
- The new feature flag `use_trace_attributes_json` (default off) now
gates the JSON columns in the traces `getColumn`, alongside the
evolution entry.
- `SelectEvolutionsForColumns` now ignores evolutions of columns the
mapper didn't return instead of erroring, so a flag-off attribute
resolves to its map even though the key carries the JSON evolution.
Part of https://github.com/SigNoz/signoz/pull/12966
#### Additional Information
- Evolution entry migration: SigNoz/signoz-otel-collector#928
- Original QB PR: #4781
#### Description
- Reverts #13014 and #13015. The resource middleware goes back to
reading body-derived resource ids with `BodyJSONPath` / `BodyJSONArray`
over the raw body, and handlers decode their own request bodies again.
- Authz should not own request decoding; that ownership stays with the
handlers.
#### Additional Information
- Contributes to: https://github.com/SigNoz/keystone-pod/issues/37
#### Description
- Follows #13014. Moves the remaining body-derived resource ids (gateway
limits, zeus hosts, cloud integration check-ins, auth domains, query
range) off gjson and onto the decoded request, with the handlers reading
the same value. Part of SigNoz/keystone-pod#37.
- Removes `BodyJSONPath`, `BodyJSONArray`, and
`ExtractorContext.RequestBody`.
#### Issues closed by this PR
- Closes: https://github.com/SigNoz/keystone-pod/issues/37
<!--A few plain bullets saying what changed and why, for a reviewer
skimming it - not a wall of text, not a restatement of the diff, not
generated boilerplate.-->
#### Description
Materialized existence checks now render as an explicit comparison
instead of a bare bool column. Results are unchanged; only skip-index
usage improves.
```sql
-- before
WHERE `attribute_string_gen_ai$$request$$model_exists`
OR `attribute_string_gen_ai$$provider$$name` = 'anthropic'
-- after
WHERE `attribute_string_gen_ai$$request$$model_exists` = true
OR `attribute_string_gen_ai$$provider$$name` = 'anthropic'
```
<details>
<summary>EXPLAIN indexes = 1 (trace-matching phase, 123M
spans)</summary>
Before: bare `col_exists`
```
Name: idx_gen_ai_span_exists
Granules: 15193/15193
Name: <Combined skip indexes>
Granules: 15193/15193
```
After: `col_exists = true`
```
Name: idx_gen_ai_span_exists
Granules: 15193/15193
Name: <Combined skip indexes>
Granules: 488/15193
```
</details>
----
- ClickHouse can use a different skip index for each side of an OR and
union the results, but it can't when one side is a bare bool column.
Comparing with `= true` fixes that.
- This shape comes from the AI explorer trace list with a span filter: a
trace qualifies when it has a gen_ai span *and* a span matching the
filter (possibly different spans), so the WHERE is `(gen_ai gate) OR
<filter>` followed by a HAVING.
- Needs the gen_ai materialized columns and `idx_gen_ai_span_exists`
from SigNoz/signoz-otel-collector#929; without them there's no index to
combine.
<!--Reference issues using `Closes #issue-number` to enable automatic
closure on merge. -->
#### Issues closed by this PR
Part of https://github.com/SigNoz/nerve-pod/issues/282
<!--Anything reviewers should keep in mind while reviewing -->
#### Additional Information
- Benchmarked the AI trace list filtered on `gen_ai.provider.name`
against a 123M-span table (direct I/O, caches off): from ~30M spans in
the window, latency drops 16–17% and CPU 35–38%, with ~25x fewer rows
read (123M spans: 510 → 427 ms, 1.5 → 0.9 sCPU). The saved time and CPU
keep growing with span count, so larger windows save more.
- Single-condition filters (`gen_ai.request.model EXISTS` in dashboard
panels, the AND-ed gate in AI aggregations) already pruned with the bare
form; no change there.
#### Description
- A logs filter on a bare key that lives in **both** resource and
another context (body or scope) ANDed the two: the resource candidate
built the `__resource_filter` fingerprint CTE while the other candidate
landed as a required main-query term, so the query matched almost
nothing.
- `ResolveLogicalFields` only preferred resource over `attribute`.
Generalized it to prefer resource over **any** other context (attribute,
body, scope, …); other contexts stay reachable via their qualified names
(e.g. `body.service.name`).
#### Issues closed by this PR
ClosesSigNoz/engineering-pod#6086
Part of https://github.com/SigNoz/platform-pod/issues/3158
#### Additional Information
Generalized rather than special-casing body/scope, since any future
context would hit the same fingerprint-CTE trap.
The semantic convention family as first call citizen revealed that the
current state of the query builder needs a bit refactoring for long term
maintenance.
The `FieldMapper` and `ConditionBuilder` are now one abstraction
`Storage`.
A storage now answers
- what the compiler cannot know i.e one read per field key (the bare
SQL, the membership present or absent, what an absent row reads, and
whether the read keeps its type or filters only).
- the fallback for a key metadata does not report
- its traits
- and one Condition compilation part.
And we introduce a new type to use in the system, `Resolved`
```
// Resolved is what resolution produces for one key: its meanings, and how
// they came to be. It is the only thing the compilers receive. Compile it
// with the operator and value it was resolved with.
type Resolved struct {
Key *telemetrytypes.TelemetryFieldKey
Fields []*telemetrytypes.LogicalField
// FromFallback: the fields came from the storage's fallback, not from
// metadata matches.
FromFallback bool
// Ambiguous: the matches held several interpretations.
Ambiguous bool
// Skipped: the storage contributes nothing for this key.
Skipped bool
Warnings []string
}
```
The prepared SQL has no changes, where it changed, it specifically made
the expression better by removing the redundant part.
- The prepared SQL remains identical with this refactoring
- No changes to integration tests
Assisted-by: Claude Fable 5.1
#### Description
- Every user-controlled field name that reaches generated SQL goes
through the new `pkg/clickhousesql` package (`Identifier`,
`StringLiteral`, `Literal`, `LikePattern`): map reads and `mapContains`,
JSON sub-column paths and the JSON body access plan, labels, fingerprint
labels, materialized column names, select aliases, group-by and order-by
references, the legacy string-body JSONPath, and the raw SQL in the
trace funnel, trace detail and infra monitoring modules. Filter
expressions built from request or telemetry values use
`querybuilder.FilterStringLiteral`. The same package now also renders
dashboard variable values in the querier, LIKE patterns in the metadata
store and label lists in the PromQL transpiler, which each had their own
escaping.
- A `$` followed by a digit, `{` or `?` is written as `\x24`, which
ClickHouse decodes in identifiers and literals. Those are the forms the
tools react to: go-sqlbuilder resolves `$0` in a compiled fragment to
its own WHERE clause and recurses until the stack overflows, and
clickhouse-go rejects a query mixing `$<digits>` with `?` arguments. Any
other `$` stays literal, so materialized column names keep their `$$`
and render exactly as before; a key like `http.2xx` becomes ``
`attribute_string_http$\x242xx` `` instead of failing in the driver.
- Compiled sqlbuilder fragments (Select, GroupBy, OrderBy, raw Where
text) are wrapped with `sqlbuilder.Escape`; the metrics builder escapes
its compiled time-series subquery, which is compiled a second time when
joined.
- The raw statement validator (`ErrIfStatementIsNotValid`,
`LogIfStatementIsNotValid`) moves from
`pkg/querybuilder/clickhouse_sql.go` to
`pkg/clickhousesql/statement.go`. Its `Code*` identifiers drop the
`ClickHouseSQL` prefix; the code strings are unchanged.
- Unit tests round-trip the helpers over hostile names and drive them
through the modules' raw SQL;
`tests/integration/tests/queriercommon/08_field_name_quoting.py` and
`querier_json_body/07_field_name_quoting.py` query such names through
the logs, traces and metrics builders against a real ClickHouse.
#### Additional Information
- `docs/contributing/go/clickhousesql.md` documents the quoting
functions, where `sqlbuilder.Escape` belongs, the `$` rule and the
statement validator; `.claude/rules/go-contrib.md` points at it.
- `pkg/clickhousesql` is a leaf package so `telemetrytypes` (JSON access
plan) and `querybuilder` share one implementation without a cycle.
- For names without special characters the generated SQL is byte
identical.
- Not covered here: the legacy v3/v4 query_range builders and the
`pkg/query-service/utils` quoting helpers (`QuoteEscapedString`,
`QuoteEscapedStringForContains`, `ClickHouseFormattedValue`,
`AddBackTickToFormatTag`), the collector's `JSONSubColumnIndexExpr`, and
aggregation arguments naming a key that contains a backtick (rejected by
the SQL parser, a 500 as before).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
#### Description
- Adds `GET|POST /prometheus/api/v1/query_range` and
`/prometheus/api/v1/query` (`pkg/prometheus/promapi`), following the
Prometheus HTTP API contract: float-unix or RFC3339 times, float-seconds
or duration-string durations, the `{status, data, errorType, error,
warnings, infos}` envelope with Prometheus' status codes, and the
11,000-point cap.
- The `/prometheus` prefix works as a drop-in Prometheus base URL:
Grafana's Prometheus data source, promtool, and the PromQL compliance
tester append `/api/v1/*` to a base URL, so they can point at SigNoz
unmodified. Same layout as Mimir/Cortex.
- Wired through `signoz.Handlers` (`prometheus.Handler` interface,
constructed in `NewHandlers`) like the other domain handlers.
- Range queries serve through the `RangeExecutor` capability when the
provider has it, so a clickhousev2-serving deployment transpiles through
these endpoints too.
- New `promapiconformance` integration suite: the frozen promqltest
corpus replayed against these endpoints with `prometheus::provider:
clickhousev2` — the two paths nothing else exercises (v2 as serving
provider, and this API surface). Instant cases go through `/query` with
a real `time` parameter. The `instant-coarse` corpus variants are
skipped — they exist only to encode instant evals as coarse ranges for
the v5 API, and their transpiled coarse-step serving is already covered
and ledgered by promqlconformance's clickhousev2 leg — so this suite
asserts zero divergences with no ledger of its own.
- Purely additive: the existing `GET /api/v1/query_range` and `GET
/api/v1/query` handlers are untouched. `openapi.yml` is generated and
these mux-registered routes are outside the generator, so their
documentation is the upstream Prometheus API contract they follow.
#### Additional Information
Final slice of the clickhouseprometheusv2 stack (#12323, #12324, #12325
— merged). Legacy endpoint removal, if ever, is a separate change after
usage drains.
#### Description
- A referenced name in a trace query now resolves to a `LogicalField`
(#12499): one field, addressed by the requested spelling, backed by its
physical member keys. A semantic-convention family
(`deployment.environment.name` / `deployment.environment`) merges into
one expression with current-wins precedence; the response keeps the
requested spelling.
- `FieldMapper` gets one new method, `ExistsFor` (the per-key presence
primitive). `LogicalValueExpr` and `LogicalExistsExpr` build all family
SQL in one place from `FieldFor` and `ExistsFor`; no signal implements
family logic.
- Statement builders prefetch sibling spellings; the metadata store
stays family-blind and autocomplete stays literal. Traces and the
resource filter compile per logical field; logs, metrics, and the other
signals keep their SQL unchanged.
- The `resolve_semconv_families` feature flag (default: disabled) gates
all family behavior. With the flag off, the generated SQL is the same as
main; tests pin this. Part of #6143.
#### Additional Information
- Stack: #12441 (merged) → **#12442** → #12443 → #12444 → #12445 →
#12446 → #12447. This layer bases on main.
- Rollback: turn the flag off; stored telemetry is untouched.
### Description
Bumps `github.com/AfterShip/clickhouse-sql-parser` from v0.5.5 to
v0.5.6.
- v0.5.6 parses a parenthesized left operand of a set operator (upstream
https://github.com/AfterShip/clickhouse-sql-parser/pull/312), e.g.
`(SELECT 1) UNION ALL (SELECT 2)`.
- Moves the three now-passing parenthesized set-operation cases into the
pass table in `clickhouse_sql_test.go` as regression canaries.
- Records the outstanding `NULLS FIRST|LAST` ORDER BY gap in the
known-gap table — the parser still rejects it, so it stays tracked until
fixed upstream.
Bumps `clickhouse-sql-parser` to v0.5.5, fixes the false rejection that
was left over once it landed, and closes three holes in the same
validator that the first two changes brought to light.
## The bump
**Reserved keywords as expression operands**
([#305](https://github.com/AfterShip/clickhouse-sql-parser/pull/305)).
`interval` was fixed in v0.5.4, but the same defect affected 36 other
keywords once the column appeared as an operand rather than bare.
Sweeping 94 candidates against ClickHouse 26.8.1.337, only `on` still
rejects — and ClickHouse runs that too. This one was live: `sum(limit)`
on a metric label.
**Panic on an unparseable `DEFAULT` expression**
([#306](https://github.com/AfterShip/clickhouse-sql-parser/pull/306)).
Both known cases return a parse error now instead of dereferencing nil.
The `recover` in `ErrIfStatementIsNotValid` stays — it guards the next
one of these, not these two.
[#307](https://github.com/AfterShip/clickhouse-sql-parser/pull/307) also
allows `CAST` in a table function's argument list.
## Table functions are only table functions in a table position
The parser types a call inside a table function's argument list as a
`TableFunctionExpr` as well, so the generator allow list only ever
cleared a generator whose argument was a literal. Every real dashboard
computes its row count — `numbers(greatest(1, intDiv(end_ns - start_ns,
step_ns) + 1))` — and every one was refused, on `intDiv` rather than on
`numbers`.
`TableExpr.Expr` is the only table position a SELECT can reach, so the
allow list asks that instead. Of the four places the parser builds a
`TableFunctionExpr`, two are `CREATE TABLE` paths rejected as
not-a-SELECT before the walk starts, one is `parseTableArgPrimaryExpr`,
and one is the `FROM`/`JOIN` path that wraps into a `TableExpr`.
## Three holes that were already open
Skipping argument position is only safe if nothing there can read, and
that turned out not to be true — not because of this change, but
independently of it.
**Reading functions.** `file` is both a table function and a scalar
function, and the validator never inspected scalar calls at all. On
`main` today, `SELECT file('/etc/passwd')` is accepted and returns the
file. A numeric wrapper passes ClickHouse's type check, so the row count
alone is an oracle: `numbers(length(file(x)))` yields one row per byte.
The same applies to the 42 dictionary accessors, which can be backed by
HTTP, ODBC or another database, to `catboostEvaluate`, and to the
introspection functions. All are now refused by name wherever they
appear, under `clickhouse_sql_reading_function`.
**`x IN db.table`.** ClickHouse reads this as `x IN (SELECT * FROM
db.table)`, and a qualified name on the right of `IN` parses as a
`Path`, not a `TableIdentifier` — so `SELECT * FROM t WHERE a IN
system.users` bypassed the internal-database rule entirely. Now checked,
including the `GLOBAL IN` and `NOT IN` forms.
**Quoted generator names.** The allow list matched on the formatted
name, which carries the quoting, so ``SELECT * FROM `numbers`(31)`` was
refused. It now reads the identifier the way the internal-database
branch already did.
## Effect
Replaying 72 distinct shapes of production `clickhouse_sql` that the
validator currently rejects: **64 pass, up from 59 on v0.5.4**. Two came
from the bump, three from the table-position change, and those three are
379 of the 1390 sampled occurrences. The three new rules add no false
positives to the corpus.
Of the eight left, four are correct rejections (`system` reads, `SHOW
TABLES`), one is a dashboard variable rendering as the literal `<no
value>`, one is SQL ClickHouse also rejects, and two are an open
upstream gap.
## Tests
`TestErrIfStatementIsNotValid_ShouldPassButFails` is back, holding what
remains: three forms of a parenthesised left operand of a set operator,
and `on` as a column name. It also stopped panicking — `errors.Asc`
dereferences the error it is given, so a case starting to pass took the
suite out with a SIGSEGV instead of reporting. Both refusal tables now
share one harness, bounded by the same timeout the passing table uses.
Known gap: no input is currently known to panic the parser, so the
`recover` has no test exercising it.
A numeric coercion yields seconds since epoch, so max(timestamp) came back
as 1758113657.04 and the third-party APIs "Last Seen" rendered as January
1970. Both rules key on the physical column type, since no FieldDataType
denotes a timestamp.
- a time column is never coerced: it reaches the aggregate in its native
type and the driver returns a time.Time
- a bare `timestamp` resolves to the intrinsic column alone; a same-named
numeric attribute no longer joins the candidate union, where the mixed
branches failed with "no supertype for DateTime64, Float64"
- the exists guard leaves the String bucket: `timestamp <> ''` becomes a
typed epoch-zero comparison, verified equivalent on ClickHouse 25.5
including a non-UTC server timezone
- max/min/quantile/count keep working; sum/avg and the rate_* variants now
fail at the database instead of returning a rescaled number, pinned by
integration tests until they are rejected up front
- tests/fixtures/traces.py wrote kind_string as the enum member name
("SPAN_KIND_CLIENT"), so any feature filtering on it matched nothing; it
now maps the six kinds to the exporter's form ("Client")
- Last Seen assertions pin the encoding and the exact instant, rather than
accepting either an RFC 3339 string or epoch millis
- drop the /overview/domain integration tests and the fixture surface that
served only that route; the endpoint is no longer used
- Replace the search() stub: the logs condition builder fans a case-insensitive
match of the term across every searchable column (log columns, body/body_v2,
attribute + resource maps).
- Grammar: searchCall takes a valueList (parser regenerated), so the scoped form
search('term', body, resource) parses and narrows the fan-out to those contexts.
- The visitor emits FilterOperatorSearch and flags the statement scan-heavy; the
statement builder attaches a CostGuard the querier enforces via EXPLAIN ESTIMATE
against a per-shard budget, cumulative across buckets on the window-list path.
- body_v2 gets its own, lower budget (search_max_scan_rows_json_body): toString()
rebuilds every document and no skip index prunes.
- Logs-only: other signals reject search().
- Unit tests for the per-context fan-out SQL and both budgets; integration suites
run the same matrix over the legacy body and over body_v2.
* feat(querybuilder): allow row-generator table functions
The table-function rule refuses everything, which costs a false positive
on queries that use numbers() or generateSeries() to build a dense
interval axis to join a sparse series against. Neither reads through
anything: they compute their rows from their arguments, open no file or
socket, reach no other host, and name no table, database or dictionary.
Everything else stays refused. Most table functions read through
something, and merge('system', '.*'), remote() and cluster() reach the
internal databases without producing a TableIdentifier for the database
rule to catch, so this rule is all that sees them. generateRandom is pure
but streams rows its arguments do not bound, so it stays out too.
Arguments are visited before the table function itself, which is what
keeps an allowed generator from being usable as a wrapper to smuggle a
read through.
* test(querybuilder): address review on the generator allow list
Names the allowed table functions in the rejection error, collapses the
added comments, and covers the generators in JOIN, CTE, subquery and
UNION position alongside the reads they must not be usable to smuggle.
* refactor(querybuilder): derive the allowed table function message from the map
Values the map by the spelling to name back to the caller, so the message
comes off the same data the lookup uses and cannot drift from it. Drops
the test that existed only to catch that drift.
* chore(deps): bump clickhouse-sql-parser to v0.5.4
v0.5.4 carries the four SigNoz-reported grammar fixes: unquoted `interval`
as a column name, GLOBAL before a join type and GLOBAL NOT IN, arithmetic
inside table function arguments, and the SQL-standard keyword-argument
forms of trim/substring/overlay.
Replaying the saved production corpus takes the v5 clickhouse_sql rejection
rate from 11/198 to 1/198, and the one remaining rejection is a real rule
hit rather than a parser gap. TestErrIfStatementIsNotValid_ShouldPassButFails
therefore has nothing left to hold; its queries move to _Pass as regression
canaries.
* test(querybuilder): bound how long a valid statement may take to parse
Telling an INTERVAL operator from a column named interval needs
backtracking, and v0.5.4 keeps that from going exponential by remembering
the offsets it has already failed at. Losing that memoisation would hang
the parser rather than fail it, so a 30-term case covers the shape and the
loop bounds how long any valid statement may take.
* test(querybuilder): address review on the parser bump
Collapse the added comments to fewer lines, and fail the timeout branch
with assert rather than t.Fatal so a case that does hang reports alongside
the rest.
* test(querybuilder): drop the stale signed-literal comment
* test(querybuilder): pin the numbers table function false positive
Restores TestErrIfStatementIsNotValid_ShouldPassButFails around the one
rejection the saved production corpus still produces on v0.5.4. numbers()
generates rows rather than reading through anything, so this is the
blanket table-function rule being stricter than the threat it exists for
rather than a parser gap, and the table now asserts the error code.
* test(querybuilder): trim the numbers case comment and query
Collapses the query to a single line to match the other cases in the file
and swaps the tenant's metric name for a placeholder.
* fix(querybuilder): give each clickhouse sql refusal its own error code
Every refusal used CodeInvalidInput, so the only way to tell why a statement was
refused was to match on the message. Counting the failures on a live deployment
meant regexing the parser's text out of the log, and grouping by cause meant
knowing which of five different messages belonged to the same underlying gap.
One code per reason instead, which arrives as exception.code on the warning that
LogIfStatementIsNotValid writes, so the failures can be grouped and alerted on
directly.
Pin the expected code on each failing case, so the mapping is checked rather
than assumed.
* chore(querybuilder): drop the comment above the error codes
The names say it.
* fix(querybuilder): separate the parser panic code from the parse failure
A panic is a parser defect worth alerting on; a parse failure is a grammar gap
that shows up in ordinary traffic. Sharing one code lumps the two together.
* chore(deps): bump clickhouse-sql-parser to v0.5.3
v0.5.3 stops lexing a signed number after a closing bracket as one literal, so
`(now()-1)` and `arr[1]-1` parse where they used to fail. That was the third of
the three gaps the validator trips over on real dashboard SQL, and the only one
where the statement had to be rewritten to be accepted.
Move the two queries it covers out of the failing set, which is what that test
exists to prompt. On the production corpus the rejection rate goes from 6.1% to
5.6% of v5 query shapes; the remaining two gaps, `interval` used as a column name
and the standard trim(BOTH x FROM y) syntax, are still open upstream.
* chore(querybuilder): link the upstream issue on the signed literal cases
* fix(querier): restrict user-authored clickhouse sql to read-only selects
Validate every user-authored ClickHouse statement before it runs: exactly one
statement, SELECT only, no table functions and no readonly override. Validation
runs on the rendered statement, since substituted variable values are user input
too.
Apply it to both entry points that reach the telemetry store with user SQL, the
clickhouse_sql query type in query_range and the dashboard variables query, and
replace the substring blacklist in the latter, which ran before substitution.
Also run these statements with readonly = 2 as a backstop. The connection is
shared with write paths, so it is opt-in per query and writers never set it.
* fix(querier): address review on clickhouse sql validation
Rename the validator to ValidateReadOnlySelect and collapse the multi-line error
constructions onto single lines.
Carry the parse failure in the message rather than as a wrapped cause. Neither
renderer surfaces the cause: render.Error reads only the message off the base
error, and RespondError reads Error(), which returns the cause alone and drops
the message. Both now show the same text with a 400.
Add unit tests covering statement kinds, table functions nested in joins, CTEs,
subqueries and unions, and readonly overrides.
* fix(querier): reject internal databases in user-authored clickhouse sql
Reading system or information_schema exposes grants, users and server metadata,
and ClickHouse read-only mode does not prevent it, so reject any table reference
into them.
Update the dashboard variables test for the new rejection message and cover both
a statement smuggled through a variable value and a system table read.
* fix(querier): reject unterminated block comments before parsing
The parser loops forever on an unterminated block comment, so validating a query
containing one would spin a request goroutine at full CPU instead of rejecting
it. Detect it up front, skipping comment markers that sit inside string literals.
Cover the parser panicking on a settings clause with no value in its own test,
which asserts the input still panics the parser so the case cannot quietly stop
exercising the recover.
* fix(querier): drop the block comment guard now the parser handles it
The parser no longer loops on an unterminated block comment, so the hand-rolled
scan that worked around it is redundant, and it was the riskier of the two: it
duplicated the lexer's handling of string literals and could have rejected a
legitimate query. Leave the parser as the single source of truth.
The panic case likewise no longer panics, so its test can no longer cover the
recover and is removed. The recover stays as insurance, since this reaches
user-authored SQL and the parser has regressed this way before.
Cover INTERSECT and EXCEPT, which the parser only started accepting in this
version and which reach a second query through the same rules.
* fix(querier): log clickhouse sql validation failures instead of rejecting
The parser's grammar has gaps against SQL that ClickHouse accepts, so rejecting
whatever it cannot read would break working dashboards and alerts. Sampling a
week of production clickhouse_sql found roughly one query in eighteen tripping
one of three gaps, none of them reading anything a telemetry query should not.
Log the failure with the rendered query instead, so the gaps can be told apart
from statements that genuinely break the rules before anything is enforced.
Wrap the parse error rather than formatting it in, so a caller can recover the
parser's *ParseError and read the position off it. The tests use that to pin the
construct each gap stops at, alongside the rewrite that the parser does accept.
* chore(querier): tighten the clickhouse sql validation comments
Condense the comments to single lines, move the recover note inside the defer it
explains, separate the visitor cases, and shorten the log message. Report how
many statements were found when rejecting a multi-statement query.
* fix(querier): log invalid clickhouse sql on the dashboard variables path too
The two callers disagreed: query_range logged and carried on, while dashboard
variables rejected. Given the parser trips on roughly one real query in eighteen,
that path could refuse a legitimate variable query, and being a rejection it left
nothing behind to show it had happened.
Both now go through LogIfStatementIsNotValid. prepareQuery is back to rendering
and nothing else, so the cases that expected it to police the statement move to
where the rules actually live.
* fix(querier): drop the read-only clickhouse session
Setting readonly on the session was defence in depth for a validator that now
only logs, so it guarded nothing while adding a context key and a settings branch
to a connection that is shared with every write path.
Restore the dashboard variables handler to what it was and call the validation
alongside, rather than reworking the rendering to accommodate it.
* chore(querybuilder): drop the redundant suffix from the failing case names
The test they sit in is already named _Fail.
* chore(deps): bump clickhouse-sql-parser to v0.5.2
The parser hangs in an infinite loop on an unterminated block comment and panics
on a settings clause with no value, both reachable from user-authored SQL. v0.5.2
fixes those, along with Accept and Walk missing AST children, which anything
walking the tree to inspect a query depends on.
String() on Expr is gone in favour of FormatSQL, so the callers that rendered a
node back to SQL now go through the Format helper, which formats compactly the
way String() did.
* test(querierlogs): cover the rate aggregation family
rate and rate_sum divide by the query window rather than the step, and nothing
exercised that path, so a change to how the aggregation expression is rendered
would have gone unnoticed. Assert both against the inserted logs.
The window is now passed to the request helper instead of relying on its default,
since these are the only assertions in the file that depend on it.
* test(querierlogs): match the response rounding for rate expectations
Scalar responses round floats to three significant figures below one, and the
rates are the only values in this test that do not divide evenly, so compare
against the rounded form rather than the exact quotient.
* test(queriertraces): cover the rate aggregation family
Traces derives the scalar rate window separately from logs, so a divergence
between the two would go unnoticed with only the logs case. rate_sum here is
taken over the intrinsic duration column, which also puts it on the other side
of the response rounding boundary from the logs test.
Add a search('needle') function to the filter grammar, recognized but not
yet executed - using it returns "search is not yet supported" instead of a
syntax error. Alert links, dashboard variables, and the contradiction check
pass it through untouched. `search` is now reserved, like has/hasAny.
* feat(authz): provision telemetry roles with plaintext selectors and hashed tuples
Accept a user-facing telemetry selector string (<query_type>/<key>/<value>
with trailing wildcards), validate and canonicalize it, store it as-is in the
role JSON record, and hash it only at the OpenFGA boundary in Object(). Grant
and check both flow through Object() so the hashes match; the plaintext record
stays the readable source of truth for display and recreation.
Because the hash is one-way, role Update diffs at the tuple level via
DiffTuples instead of reconstructing transaction groups from stored tuples.
* refactor(authz): rename telemetry selector helpers and unexport hash
Rename grant_selector.go to selector.go, CanonicalizeTelemetryGrantSelector to
NewTelemetryGrantSelector, and CanonicalTelemetryGrantKey to NewTelemetryGrantKey.
Unexport telemetrySelectorHash since Object() is its only caller.
* fix(authz): expand telemetry ladder in the permission check API
The check API only probed the exact selector plus the full wildcard, so a
scoped grant like builder_query/service.name/* reported a concrete query as
unauthorized even though enforcement allowed it. Relocate the grant selector
ladder to telemetrytypes as NewTelemetryGrantSelectors so both enforcement and
the check API share it, and canonicalize plus fan out the ladder for telemetry
check transactions. Non-telemetry resources keep the exact-plus-wildcard probe.
* refactor(authz): move newCheckSelectors to the bottom of tuple.go
* fix(authz): reject key-scoped selectors for promql and clickhouse_sql
Only builder_query and builder_sub_query extract key-scoped selectors on the
check side, so a grant like clickhouse_sql/service.name/signoz could never
match anything. Track key-scope support per query type and reject the
<query_type>/<key>/<value> form for query types that only emit <query_type>/*.
* test(authz): use a valid promql selector in the check API test
The promql/service.name/service-a case became invalid input after key-scoped
promql/clickhouse_sql selectors were rejected, so the check endpoint returned
400 instead of 200. Use promql/* to keep exercising the different-query-type
denial with a valid selector.
* feat(authz): allow signoz.workspace.key.id as a telemetry grant key
Add signoz.workspace.key.id to the telemetry grant key allowlist so roles can
scope telemetry access by ingestion key, alongside service.name. Grant
validation and check-side extraction both pick it up through the shared
NewTelemetryGrantKey path; bare and resource.-prefixed spellings fold to the
same canonical key.
* Revert "feat(authz): allow signoz.workspace.key.id as a telemetry grant key"
This reverts commit c50d122b54.
* feat(authz): scope telemetry grants by signoz.workspace.key.id
Make signoz.workspace.key.id the sole telemetry grant key instead of
service.name, so roles scope telemetry access by ingestion key. The allowlist
is the only production change; grant validation and check-side extraction are
name-agnostic. Update the querierauthz integration suite and unit tests to the
new key.
* test(authz): update query_range_resources extractor tests to signoz.workspace.key.id
The grant-key switch left this extractor test filtering on and expecting
builder_query/service.name/... IDs, which now fall back to builder_query/*.
Rename the filters and expectations to signoz.workspace.key.id.
A span/log query that aggregates or filters a numeric intrinsic column
(e.g. duration_nano) failed with ClickHouse NO_COMMON_TYPE (386) when a
same-named, type-consistent attribute existed in the metadata: field-key
resolution unioned the intrinsic column with the attribute into a multiIf/OR
whose branches had incompatible types (UInt64 intrinsic vs Float64 attribute).
This broke /api/v1/span_percentile on affected tenants.
- fallback_expr: DataTypeCollisionHandledFieldName now handles the unspecified
data type in the projection/aggregation path (operator Unknown), coercing the
column so collision multiIf branches share a supertype. Comparison contexts
are left bare so the column index stays usable.
- telemetrytraces/condition_builder: run collision handling for duration_nano
(so a same-named string attribute is cast in comparisons) and coerce numeric
duration values to int64, keeping the intrinsic comparison bare/index-friendly
while preserving the duration-string QoL parsing.
- tests: add a type-consistent "collision" trace_noise variant; cover it in the
percentile aggregation test and add a duration_nano QoL filter regression test.
* fix: convert key not found to warnings
* fix(logs): has-family body-only errors, not-found warnings, and condition-builder cleanup
- has/hasAny/hasAll/hasToken on a non-body key now return "supports only body
JSON search" (both modes); a not-found body path warns and queries the
underlying data instead of 400 (JSON mode)
- capture the two has-family 500s (hasToken separator/whitespace needle;
hasAny/hasAll quoted int >= 2^32) as regression tests
- use telemetrytypes.NewTelemetryFieldKey for derived keys (fresh identity-only
keys, no stale resolved metadata) instead of struct copies of the field key
- rename ConditionForKeys -> ConditionFor and inline conditionsForKeys into it
across the builders; rename private conditionFor -> conditionForResolvedKey
- move the trace_noise fixture from queriertraces/conftest.py to fixtures/traces.py
* fix: convert trace_noise fixture to util
* feat(authz): enable FGA for telemetry resources on v5 query_range
Authorize /api/v5/query_range and /preview at the telemetry-resource level,
derived from the request body:
- coretypes: ResourceWithID + ResourceExtractor as the resource-level analogue
of the id extractors; NewResolvedResourceWithID/NewResolvedResourceWithError;
telemetryresource selector regex widened to query-type selectors with up to
two hashed segments (metric name, where clause) or wildcards
- telemetrytypes: QueryRangeResources maps each query to its telemetry
resource (signal/source aware: audit-logs, meter-metrics) with a hierarchical
selector id (query_type/<hash(metric)>/<hash(where)>); PrefixSelector expands
the id into the grant ladder [exact, prefix/*..., *]
- handler: generic TelemetryResourceDef fans out an injected ResourceExtractor;
fails closed when extraction errors or resolves nothing
- audit: log and skip resolved resources that carry a resolution error
- querier routes: ViewAccess -> CheckResources with telemetry read scopes;
substitute_vars stays ViewAccess (no telemetry access)
- sqlmigration 099: backfill telemetry read tuples for existing orgs
(admin: logs/traces/metrics/audit-logs/meter-metrics; editor/viewer:
logs/traces/metrics)
* feat(authz): widen telemetry selector segments to 128 bits
64-bit truncation permits chosen-collision attacks at ~2^32 work; 128 bits
pushes this to 2^64. No hashed selector is persisted yet, so the change is
free.
* chore(docs): regenerate openapi spec with telemetry read scopes
* feat(telemetry): add where clause visitor
* refactor(telemetry): restructure normalizer file and quote bare values
* feat(authz): gate v5 query_range on service.name telemetry selectors
* feat(authz): encode telemetry grants as query-type qualified atom selectors
* feat(authz): move telemetry grant key to plaintext selector segment
* feat(authz): use escaped plaintext telemetry selectors with mechanical ladder
* revert(authz): restore transaction group diff in role update
* test(authz): add querierauthz integration suite for telemetry query_range gating
* test(authz): seed logs so service.name resolves in allowed querierauthz cases
* feat(authz): backfill telemetry read tuples for existing orgs
* chore(authz): reword empty composite query error message
* feat(authz): add meter metrics and audit logs to clickhouse sql
* Revert "feat(authz): add meter metrics and audit logs to clickhouse sql"
This reverts commit c9d870e0ee.
* feat(authz): grant meter-metrics to editor/viewer, keep clickhouse admin-only
* feat(authz): remove the audit logs from clickhouse check altogether until it's introduced
Repairs 10 broken signoz.io/docs links (5 hard 404s + 5 dead anchors)
that survived the frontend-only sweep in #11319 because they live in
the Go backend and two frontend files it did not cover.
- infra-monitoring readiness checks: drop the removed `user-guides/`
path segment and remap to the current hostmetrics/k8s-metrics anchors
- querybuilder / telemetrylogs search-troubleshooting errors: point to
the reworded Q&A anchors (update matching test assertion)
- alert generatorURL fallback: `alerts-management/#generator-url` ->
`alerts/` (anchor removed in docs restructure)
- missing-spans banner: -> traces-management troubleshooting FAQ anchor
- agent-skills install link: `#installation` -> `#install-the-plugin`
Every changed URL verified live (200 + anchor present).
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore: mark required and nullable in json tag, renamed methods, and added more functionality
* fix: unit test
* fix: cast clickhouse exceptions
* fix: go mod tidy
* fix: telemetrystore now returns explicit base errors
* fix: typo
* fix: added all changes
* fix: added nil check
* fix: update test files
* fix: addressed comments
* fix: change errors and suggestions to be non-nullable
* feat: extend error responses with new error struct
* fix: enriched error for dashboard api
* fix: merge issues
* fix: reverted dashboards changes and add for cloud integrations
* fix: delete file
* fix: add back file
* fix: added a helper
* fix: removed invlaid referencess
* fix: generate openapi
* fix: keeping additional along with suggestion
* Revert "fix: keeping additional along with suggestion"
This reverts commit be30e2ffd2.
* fix: added suggestions per additonal error
* fix: generate openapi
* fix: remove valid references
* fix: removeg valid references for select and group by and only did you mean is kept
* fix: unit test
* fix: use binding for deconding for both ee and community
* fix: trim down suggestions methods
* fix: added renamed methods and moved stuff around
* fix: typo
* fix: removed json decoder
* fix: added empty check
* fix: retain addtional
* fix: reverted re-structing of file
* fix: fts warning miss in direct text search
* fix: comments
* test: added one more test variation
* ci: go lint
* fix: fts warning update
* fix: integration tests
* fix: go test and fmtlint
* chore: add json enabled as feature flag for FE
* fix: still using global bool
* feat: flagger integration in flow
* fix: flagger threaded into tests
* test: removed nil checks
* fix: minor changes
* chore: rename field
* chore: remove querybuilder helper
* fix: unit tests
* fix: correct env var
* fix: lint fix
* fix: lint
* chore: replace flag
* fix: tons of changes
* chore: remove redundent comparison
* ci: tests fixed
* fix: upgraded collector version
* fix: qbtoexpr tests
* fix: go sum
* chore: upgrade collector version v0.144.3-rc.4
* fix: tests
* ci: test fix
* revert: remove db binaries
* test: selectField tests added
* fix: added safeguards in plan generation
* fix: name changed to field_map
* fix: json access plan remval of AvailableTypes
* fix: invalid index usage on terminal condition
* fix: branches should tell missing array types
* fix: comment removed
* fix: issue with FuzzyMatching and API failing
* fix: int64 mapping
* ci: test and lint fix
* fix: test VisitKey
* test: running test for sku
* fix: buildFieldForJSON works
* fix: few minor changes
* fix: refactor tag vs field_key table
* fix: minor changes based on review
* revert: minor variable change
* fix: added more membership testcases
* revert: minor var names reverted
* ci: tests aligned
* fix: indexed expressions
* refactor: move resourcefilter to pkg/telemetryresourcefilter
Move pkg/querybuilder/resourcefilter to pkg/telemetryresourcefilter
to align with the existing telemetry package naming convention
(telemetrylogs, telemetrytraces, telemetrymetrics, telemetrymeter).
The resource filter is a statement builder, not a query builder utility.
* refactor: internalize resource filter construction in statement builders
Each telemetry statement builder (logs, traces) now creates its own
resource filter internally instead of receiving it as an injected
dependency. This makes it impossible to wire the wrong resource table
and simplifies the provider.
Delete telemetryresourcefilter/tables.go — each telemetry package now
owns its resource table constant (LogsResourceV2TableName in
telemetrylogs, TracesResourceV3TableName in telemetrytraces).
* refactor: create field mapper and condition builder inside resource filter New
Remove fieldMapper and conditionBuilder params from
telemetryresourcefilter.New — they are always the same
(NewFieldMapper + NewConditionBuilder) so create them internally.
* fix: added validations for having expression
* fix: added extra validation and unit tests
* fix: added antlr based parsing for validation
* fix: added more unit tests
* fix: removed validation on having in range request validations
* fix: generated lexer files and added more unit tests
* fix: edge cases
* fix: added cmnd to scripts for generating lexer
* fix: use std libg sorting instead of selection sort
* fix: support implicit and
* fix: allow bare not in expression
* fix: added suggestion for having expression
* fix: typo
* fix: added more unit tests, handle white space difference in aggregation exp and having exp
* fix: added support for in and not, updated errors
* fix: added support for brackets list
* fix: lint error
* fix: handle non spaced expression
---------
Co-authored-by: Srikanth Chekuri <srikanth.chekuri92@gmail.com>
* fix: handle empty not() expression
* fix: handle more cases
* fix: short circuit conditions and updated unit tests
* fix: revert commented code
* fix: added more unit tests
* fix: added integration tests
* fix: make py-lint and make py-fmt
* fix: moved from traces to logs for testing full text search
* fix: simplify code
* fix: added more unit tests
* fix: addressed comments
* fix: update comment
Co-authored-by: Srikanth Chekuri <srikanth.chekuri92@gmail.com>
* fix: update unit test
Co-authored-by: Srikanth Chekuri <srikanth.chekuri92@gmail.com>
* fix: update unit test
Co-authored-by: Srikanth Chekuri <srikanth.chekuri92@gmail.com>
* fix: instead of using true, using a skip literal
* fix: unit test
* fix: update integration test
* fix: update unit for relevance
* fix: lint error
* fix: added a new literal for error condition, added more unit tests
* fix: merge issues
* fix: inline comments
* fix: update unit tests merging from main
* fix: make py-fmt and make py-lint
* fix: type handling
---------
Co-authored-by: Srikanth Chekuri <srikanth.chekuri92@gmail.com>
* feat(instrumentation): add OTel exception semantic convention log handler
Add a loghandler.Wrapper that enriches error log records with OpenTelemetry
exception semantic convention attributes (exception.type, exception.code,
exception.message, exception.stacktrace).
- Add errors.Attr() helper for standardized error logging under "exception" key
- Add exception log handler that replaces raw error attrs with structured group
- Wire exception handler into the instrumentation SDK logger chain
- Remove LogValue() from errors.base as the handler now owns structuring
* refactor: replace "error", err with errors.Attr(err) across codebase
Migrate all slog error logging from ad-hoc "error", err key-value pairs
to the standardized errors.Attr(err) helper, enabling the exception log
handler to enrich these logs with OTel semantic convention attributes.
* refactor: enforce attr-only slog style across codebase
Change sloglint from kv-only to attr-only, requiring all slog calls to
use typed attributes (slog.String, slog.Any, etc.) instead of key-value
pairs. Convert all existing kv-style slog calls in non-excluded paths.
* refactor: tighten slog.Any to specific types and standardize error attrs
- Replace slog.Any with slog.String for string values (action, key, where_clause)
- Replace slog.Any with slog.Uint64 for uint64 values (start, end, step, etc.)
- Replace slog.Any("err", err) with errors.Attr(err) in dispatcher and segment analytics
- Replace slog.Any("error", ctx.Err()) with errors.Attr in factory registry
* fix(instrumentation): use Unwrapb message for exception.message
Use the explicit error message (m) from Unwrapb instead of
foundErr.Error(), which resolves to the inner cause's message
for wrapped errors.
* feat(errors): capture stacktrace at error creation time
Store program counters ([]uintptr) in base errors at creation time
using runtime.Callers, inspired by thanos-io/thanos/pkg/errors. The
exception log handler reads the stacktrace from the error instead of
capturing at log time, showing where the error originated.
* fix(instrumentation): apply default log wrappers uniformly in NewLogger
Move correlation, filtering, and exception wrappers into NewLogger so
all call sites (including CLI loggers in cmd/) get them automatically.
* refactor(instrumentation): remove variadic wrappers from NewLogger
NewLogger no longer accepts arbitrary wrappers. The core wrappers
(correlation, filtering, exception) are hardcoded, preventing callers
from accidentally duplicating behavior.
* refactor: migrate remaining "error", <var> to errors.Attr across legacy paths
Replace all remaining "error", <variable> key-value pairs with
errors.Attr(<variable>) in pkg/query-service/ and ee/query-service/
paths that were missed in the initial migration due to non-standard
variable names (res.Err, filterErr, apiErrorObj.Err, etc).
* refactor(instrumentation): use flat exception.* keys instead of nested group
Use flat keys (exception.type, exception.code, exception.message,
exception.stacktrace) instead of a nested slog.Group in the exception
log handler.