Compare commits

..

4 Commits

Author SHA1 Message Date
Tushar Vats
b2b6410e14 perf(logs): let attribute filters use the mapValues bloom filters
attributes_string and attributes_number each carry a bloom filter over mapValues, but the
subscript a filter reads matches no index expression, so a filter on a key every row carries -
a status code, a client address - reads every granule. The mapKeys index prunes only when the
key itself is rare. Both cases below AND in a predicate over mapValues that the original filter
already implies. On 1M rows, 1000001 rows read down to 8192, returning the same rows.

- A numeric equality carries the membership it implies. conditionForKey pairs a positive
  comparison with mapContains, so the key is present and the value compared is one of the map's
  own - which is what makes the assertion hold for the zero a subscript returns for an absent
  key. A non-numeric value means the collision handler compared the column as text, where an
  Array(Float64) membership check has no supertype, so it carries no predicate.
- A case-insensitive match reaches the raw-valued index only for a pattern holding no ASCII
  letter, LOWER being the identity on those bytes; a letter breaks the implication, since an `a`
  in the pattern may have come from an `A` in the value. That covers the values case folding
  cannot alter: addresses, timestamps, ports, numeric ids.

Both read a column the query already reads - a Map subscript decompresses keys and values
either way - so when the filter does not prune, bytes read are unchanged.
2026-08-19 18:49:24 +05:30
Tushar Vats
b7cc59d85f refactor(logs): drop the needle term from the logs condition builder
The word carried three different meanings in one package: the substring a filter asserts over
the indexed text, the value a has-family filter looks for in an array, and the single token
hasToken matches. Each is now named for what it is - literal, element, token - so the reader
does not have to infer which one a given `needle` refers to.

Renames only. `castNeedleArray` and `legacyCoerceNeedle` follow their parameters, and the
hasToken error keeps its wording, which was always about tokens.
2026-08-19 18:49:23 +05:30
Tushar Vats
79b72adb17 perf(logs): let the legacy has family use the lower(body) bloom filters
has, hasAll and hasAny over a legacy body path compare extracted JSON, which matches no index
expression, so they read every granule. They now carry what they imply over the indexed
LOWER(body), ANDed onto the extraction that still decides the row:

- the path, since the extraction yields nothing for an absent one;
- each element, since it must appear in the text. has and hasAll require every element, so each
  becomes its own predicate; hasAny requires only one, so the arms are ORed - and an element
  yielding no literal leaves that OR unassertable, which drops it whole.

Only strings imply text: the family compares at the element type it infers, so a quoted number is
still a number here and its digits need not appear in the body - the same reason `=` skips them.

The comparison and its literals are built together, so the element type and the element list are
resolved once instead of in two places.
2026-08-19 18:49:03 +05:30
Tushar Vats
99c8209e94 perf(logs): let body comparisons use the lower(body) bloom filters
logs_v2 indexes lower(body) with a token and an ngram bloom filter, but nothing a body filter
compares matches that expression, so every one of them reads all granules. Each case below ANDs
in a predicate over the indexed expression that the original filter already implies: the bloom
filters prune on it, the original still decides the row.

- `body = ?` carries the lowered comparison, and IN inherits it per arm through the `=`
  delegation. On 1M rows, 123/123 granules down to 1/123.
- A legacy body JSON filter compares JSON_VALUE output, so it carries literals over the raw body
  text instead: the quoted key name rides the existence assertion, the value rides the
  comparison. Literals stop at bytes a JSON encoder may rewrite, and a number carries none
  because JSONExtract reads 1.23e2 as 123.

`body LIKE` was also silently case-insensitive: it rendered LOWER(body) LIKE LOWER(?), the shape
v3/v4 chose to reach the index. It now renders a case-sensitive LIKE with that lowered comparison
beside it, so the operator means what it says and still prunes. NOT LIKE gets no companion - a
lowered one would drop rows differing from the pattern only in case, and a bloom filter cannot
prune a negation anyway.

The has family is left alone here and follows in its own change.

The integration cases assert the read funnel from the preview endpoint rather than the SQL text:
ClickHouse lists a skip index only when the predicate matches its expression, so a companion that
stops matching drops out of the funnel entirely.
2026-08-19 18:48:25 +05:30
33 changed files with 1354 additions and 1114 deletions

View File

@@ -8807,7 +8807,6 @@ components:
- span
- trace
- resource
- scope
- attribute
- body
- ""

View File

@@ -3492,7 +3492,6 @@ export enum TelemetrytypesFieldContextDTO {
span = 'span',
trace = 'trace',
resource = 'resource',
scope = 'scope',
attribute = 'attribute',
body = 'body',
'' = '',

View File

@@ -18,9 +18,10 @@ jest.mock('periscope/components/DataViewer', () => ({
DataViewer: (): JSX.Element => <div data-testid="overview-data-viewer" />,
}));
// Force v2 for these tests regardless of route.
jest.mock('../useIsLogDetailsV2', () => ({
useIsLogDetailsV2: (): boolean => true,
// The flag to be removed later
jest.mock('../constants', () => ({
...jest.requireActual('../constants'),
isLogDetailsV2: true,
}));
const mockLog: ILog = {

View File

@@ -1,3 +1,6 @@
// temporary flag to be removed with old log details code.
export const isLogDetailsV2 = true;
export const VIEW_TYPES = {
OVERVIEW: 'OVERVIEW',
JSON: 'JSON',

View File

@@ -51,12 +51,11 @@ import { ILogBody } from 'types/api/logs/log';
import { Query, TagFilter } from 'types/api/queryBuilder/queryBuilderData';
import { DataSource, StringOperators } from 'types/common/queryBuilder';
import { RESOURCE_KEYS, VIEW_TYPES, VIEWS } from './constants';
import { isLogDetailsV2, RESOURCE_KEYS, VIEW_TYPES, VIEWS } from './constants';
import { LogDetailInnerProps, LogDetailProps } from './LogDetail.interfaces';
import LogDetailsHeader from './LogDetailsHeader/LogDetailsHeader';
import { useLogNavigation } from './LogDetailsHeader/useLogNavigation';
import LogHighlights from './LogHighlights/LogHighlights';
import { useIsLogDetailsV2 } from './useIsLogDetailsV2';
import './LogDetails.styles.scss';
@@ -93,8 +92,6 @@ function LogDetailInner({
const [isEdit, setIsEdit] = useState<boolean>(false);
const { stagedQuery } = useQueryBuilder();
const isLogDetailsV2 = useIsLogDetailsV2();
// Handle clicks outside to close drawer, except on explicitly ignored regions
useEffect(() => {
const handleClickOutside = (e: MouseEvent): void => {

View File

@@ -1,9 +0,0 @@
import ROUTES from 'constants/routes';
import { useLocation } from 'react-router-dom';
// v2 is rolled out only on the logs explorer route for now; every other surface
// (dashboards, infra monitoring, etc.) keeps the v1 log details view.
export function useIsLogDetailsV2(): boolean {
const { pathname } = useLocation();
return pathname === ROUTES.LOGS_EXPLORER;
}

View File

@@ -10,7 +10,6 @@ const fieldContextToSuggestionMap: Record<
[TelemetrytypesFieldContextDTO.attribute]: 'attribute',
// no maps for the following values on suggestion context
[TelemetrytypesFieldContextDTO.trace]: undefined,
[TelemetrytypesFieldContextDTO.scope]: undefined,
[TelemetrytypesFieldContextDTO.body]: undefined,
[TelemetrytypesFieldContextDTO.metric]: undefined,
[TelemetrytypesFieldContextDTO.log]: undefined,

View File

@@ -13,7 +13,7 @@ import { ChangeViewFunctionType } from 'container/ExplorerOptions/types';
import { OptionsQuery } from 'container/OptionsMenu/types';
import { useIsDarkMode } from 'hooks/useDarkMode';
import { ChevronDown, ChevronRight, Search } from '@signozhq/icons';
import { useIsLogDetailsV2 } from 'components/LogDetail/useIsLogDetailsV2';
import { isLogDetailsV2 } from 'components/LogDetail/constants';
import { DataViewer } from 'periscope/components/DataViewer';
import { IField } from 'types/api/logs/fields';
import { ILog } from 'types/api/logs/log';
@@ -69,8 +69,6 @@ function Overview({
isListViewPanel,
});
const isLogDetailsV2 = useIsLogDetailsV2();
if (isLogDetailsV2) {
const raw = aggregateAttributesResourcesToObject(logData);
const prettyData = buildPrettyViewData(raw);

View File

@@ -56,17 +56,6 @@ func QueryStringToKeysSelectors(query string) []*telemetrytypes.FieldKeySelector
FieldDataType: key.FieldDataType,
})
}
// todo(tushar): consider reverting changes done to this method in below PR to avoid scope specific checks
// https://github.com/SigNoz/signoz/issues/11374
if key.FieldContext == telemetrytypes.FieldContextScope {
keys = append(keys, &telemetrytypes.FieldKeySelector{
Name: key.FieldContext.StringValue() + "." + key.Name,
Signal: key.Signal,
FieldContext: telemetrytypes.FieldContextUnspecified, // this allows 'scope.' prefix for keys with other context as well
FieldDataType: key.FieldDataType,
})
}
}
}

View File

@@ -72,44 +72,6 @@ func TestQueryToKeys(t *testing.T) {
},
},
},
{
query: `scope.version = '1.0.0'`,
expectedKeys: []telemetrytypes.FieldKeySelector{
{
Name: "version",
Signal: telemetrytypes.SignalUnspecified,
FieldContext: telemetrytypes.FieldContextScope,
FieldDataType: telemetrytypes.FieldDataTypeUnspecified,
},
{
Name: "scope.version",
Signal: telemetrytypes.SignalUnspecified,
FieldContext: telemetrytypes.FieldContextUnspecified,
FieldDataType: telemetrytypes.FieldDataTypeUnspecified,
},
},
},
{
// A scope attribute whose own name carries a `scope.` prefix. `scope.prefixed`
// normalizes to {prefixed, scope}; the second selector re-adds the prefix so the
// metadata fetch can target the attribute's exact key `scope.prefixed` rather than
// relying on the broad `%prefixed%` match.
query: `scope.prefixed = 'x'`,
expectedKeys: []telemetrytypes.FieldKeySelector{
{
Name: "prefixed",
Signal: telemetrytypes.SignalUnspecified,
FieldContext: telemetrytypes.FieldContextScope,
FieldDataType: telemetrytypes.FieldDataTypeUnspecified,
},
{
Name: "scope.prefixed",
Signal: telemetrytypes.SignalUnspecified,
FieldContext: telemetrytypes.FieldContextUnspecified,
FieldDataType: telemetrytypes.FieldDataTypeUnspecified,
},
},
},
}
for _, testCase := range testCases {

View File

@@ -462,8 +462,8 @@ func TestStatementBuilderListQueryResourceTests(t *testing.T) {
Limit: 10,
},
expected: qbtypes.Statement{
Query: "SELECT timestamp, id, trace_id, span_id, trace_flags, severity_text, severity_number, scope_name, scope_version, body, attributes_string, attributes_number, attributes_bool, resources_string, scope_string FROM signoz_logs.distributed_logs_v2 WHERE (JSON_VALUE(body, '$.\"status\"') = ? AND JSON_EXISTS(body, '$.\"status\"')) AND timestamp >= ? AND ts_bucket_start >= ? AND timestamp < ? AND ts_bucket_start <= ? LIMIT ?",
Args: []any{"success", "1747947419000000000", uint64(1747945619), "1747983448000000000", uint64(1747983448), 10},
Query: "SELECT timestamp, id, trace_id, span_id, trace_flags, severity_text, severity_number, scope_name, scope_version, body, attributes_string, attributes_number, attributes_bool, resources_string, scope_string FROM signoz_logs.distributed_logs_v2 WHERE ((JSON_VALUE(body, '$.\"status\"') = ? AND LOWER(body) LIKE LOWER(?)) AND (JSON_EXISTS(body, '$.\"status\"') AND LOWER(body) LIKE LOWER(?))) AND timestamp >= ? AND ts_bucket_start >= ? AND timestamp < ? AND ts_bucket_start <= ? LIMIT ?",
Args: []any{"success", "%success%", "%\"status\"%", "1747947419000000000", uint64(1747945619), "1747983448000000000", uint64(1747983448), 10},
Warnings: []string{querybuilder.NewKeyNotFoundWarning("status")},
},
expectedErr: nil,
@@ -481,8 +481,8 @@ func TestStatementBuilderListQueryResourceTests(t *testing.T) {
Limit: 10,
},
expected: qbtypes.Statement{
Query: "SELECT timestamp, id, trace_id, span_id, trace_flags, severity_text, severity_number, scope_name, scope_version, body, attributes_string, attributes_number, attributes_bool, resources_string, scope_string FROM signoz_logs.distributed_logs_v2 WHERE ((JSON_VALUE(body, '$.\"user_names\"[*]') = ?) AND JSON_EXISTS(body, '$.\"user_names\"[*]')) AND timestamp >= ? AND ts_bucket_start >= ? AND timestamp < ? AND ts_bucket_start <= ? LIMIT ?",
Args: []any{"john_doe", "1747947419000000000", uint64(1747945619), "1747983448000000000", uint64(1747983448), 10},
Query: "SELECT timestamp, id, trace_id, span_id, trace_flags, severity_text, severity_number, scope_name, scope_version, body, attributes_string, attributes_number, attributes_bool, resources_string, scope_string FROM signoz_logs.distributed_logs_v2 WHERE (((JSON_VALUE(body, '$.\"user_names\"[*]') = ? AND LOWER(body) LIKE LOWER(?))) AND (JSON_EXISTS(body, '$.\"user_names\"[*]') AND LOWER(body) LIKE LOWER(?))) AND timestamp >= ? AND ts_bucket_start >= ? AND timestamp < ? AND ts_bucket_start <= ? LIMIT ?",
Args: []any{"john_doe", "%john\\_doe%", "%\"user\\_names\"%", "1747947419000000000", uint64(1747945619), "1747983448000000000", uint64(1747983448), 10},
Warnings: []string{querybuilder.NewKeyNotFoundWarning("user_names[*]")},
},
expectedErr: nil,
@@ -498,8 +498,8 @@ func TestStatementBuilderListQueryResourceTests(t *testing.T) {
Limit: 10,
},
expected: qbtypes.Statement{
Query: "SELECT timestamp, id, trace_id, span_id, trace_flags, severity_text, severity_number, scope_name, scope_version, body, attributes_string, attributes_number, attributes_bool, resources_string, scope_string FROM signoz_logs.distributed_logs_v2 WHERE (has(JSONExtract(JSON_QUERY(body, '$.\"user_names\"[*]'), 'Array(Nullable(String))'), ?) OR ifNull((JSON_VALUE(body, '$.\"user_names\"') = ? AND JSONType(body, 'user_names') NOT IN ('Array', 'Object', 'Null')), false)) AND timestamp >= ? AND ts_bucket_start >= ? AND timestamp < ? AND ts_bucket_start <= ? LIMIT ?",
Args: []any{"john_doe", "john_doe", "1747947419000000000", uint64(1747945619), "1747983448000000000", uint64(1747983448), 10},
Query: "SELECT timestamp, id, trace_id, span_id, trace_flags, severity_text, severity_number, scope_name, scope_version, body, attributes_string, attributes_number, attributes_bool, resources_string, scope_string FROM signoz_logs.distributed_logs_v2 WHERE ((has(JSONExtract(JSON_QUERY(body, '$.\"user_names\"[*]'), 'Array(Nullable(String))'), ?) OR ifNull((JSON_VALUE(body, '$.\"user_names\"') = ? AND JSONType(body, 'user_names') NOT IN ('Array', 'Object', 'Null')), false)) AND LOWER(body) LIKE LOWER(?) AND LOWER(body) LIKE LOWER(?)) AND timestamp >= ? AND ts_bucket_start >= ? AND timestamp < ? AND ts_bucket_start <= ? LIMIT ?",
Args: []any{"john_doe", "john_doe", "%\"user\\_names\"%", "%john\\_doe%", "1747947419000000000", uint64(1747945619), "1747983448000000000", uint64(1747983448), 10},
Warnings: []string{querybuilder.NewKeyNotFoundWarning("user_names[*]")},
},
expectedErr: nil,
@@ -1011,8 +1011,8 @@ func TestStmtBuilderBodyField(t *testing.T) {
},
enableUseJSONBody: false,
expected: qbtypes.Statement{
Query: "SELECT timestamp, id, trace_id, span_id, trace_flags, severity_text, severity_number, scope_name, scope_version, body, attributes_string, attributes_number, attributes_bool, resources_string, scope_string FROM signoz_logs.distributed_logs_v2 WHERE body = ? AND timestamp >= ? AND ts_bucket_start >= ? AND timestamp < ? AND ts_bucket_start <= ? LIMIT ?",
Args: []any{"", "1747947419000000000", uint64(1747945619), "1747983448000000000", uint64(1747983448), 10},
Query: "SELECT timestamp, id, trace_id, span_id, trace_flags, severity_text, severity_number, scope_name, scope_version, body, attributes_string, attributes_number, attributes_bool, resources_string, scope_string FROM signoz_logs.distributed_logs_v2 WHERE (body = ? AND LOWER(body) = LOWER(?)) AND timestamp >= ? AND ts_bucket_start >= ? AND timestamp < ? AND ts_bucket_start <= ? LIMIT ?",
Args: []any{"", "", "1747947419000000000", uint64(1747945619), "1747983448000000000", uint64(1747983448), 10},
},
expectedErr: nil,
},

View File

@@ -163,15 +163,31 @@ func getKeySelectors(query qbtypes.QueryBuilderQuery[qbtypes.TraceAggregation])
}
for idx := range query.GroupBy {
keySelectors = append(keySelectors, keySelectorsForField(query.GroupBy[idx].TelemetryFieldKey)...)
groupBy := query.GroupBy[idx]
keySelectors = append(keySelectors, &telemetrytypes.FieldKeySelector{
Name: groupBy.Name,
Signal: telemetrytypes.SignalTraces,
FieldContext: groupBy.FieldContext,
FieldDataType: groupBy.FieldDataType,
})
}
for idx := range query.SelectFields {
keySelectors = append(keySelectors, keySelectorsForField(query.SelectFields[idx])...)
keySelectors = append(keySelectors, &telemetrytypes.FieldKeySelector{
Name: query.SelectFields[idx].Name,
Signal: telemetrytypes.SignalTraces,
FieldContext: query.SelectFields[idx].FieldContext,
FieldDataType: query.SelectFields[idx].FieldDataType,
})
}
for idx := range query.Order {
keySelectors = append(keySelectors, keySelectorsForField(query.Order[idx].Key.TelemetryFieldKey)...)
keySelectors = append(keySelectors, &telemetrytypes.FieldKeySelector{
Name: query.Order[idx].Key.Name,
Signal: telemetrytypes.SignalTraces,
FieldContext: query.Order[idx].Key.FieldContext,
FieldDataType: query.Order[idx].Key.FieldDataType,
})
}
for idx := range keySelectors {
@@ -182,26 +198,6 @@ func getKeySelectors(query qbtypes.QueryBuilderQuery[qbtypes.TraceAggregation])
return keySelectors
}
func keySelectorsForField(key telemetrytypes.TelemetryFieldKey) []*telemetrytypes.FieldKeySelector {
selectors := []*telemetrytypes.FieldKeySelector{
{
Name: key.Name,
Signal: telemetrytypes.SignalTraces,
FieldContext: key.FieldContext,
FieldDataType: key.FieldDataType,
},
}
if key.FieldContext != telemetrytypes.FieldContextUnspecified {
selectors = append(selectors, &telemetrytypes.FieldKeySelector{
Name: key.FieldContext.StringValue() + "." + key.Name,
Signal: telemetrytypes.SignalTraces,
FieldContext: telemetrytypes.FieldContextUnspecified,
FieldDataType: key.FieldDataType,
})
}
return selectors
}
// mergeDeprecatedTraceKeys prepends deprecated intrinsic/calculated trace field
// definitions to the keys map. We do this during statement building, not at
// metadata fetch time, because:
@@ -273,14 +269,20 @@ func adjustTraceKey(key *telemetrytypes.TelemetryFieldKey, keys map[string][]*te
For example: trace_id (intrinsic), response_status_code (calculated).
*/
// Resolve against the context-qualified name first, then the bare name since that can be instrinsic field e.g. scope.name.
var isIntrinsicOrCalculatedField bool
var intrinsicOrCalculatedField telemetrytypes.TelemetryFieldKey
if key.FieldContext != telemetrytypes.FieldContextUnspecified {
intrinsicOrCalculatedField, isIntrinsicOrCalculatedField = lookupIntrinsicOrCalculatedField(key.FieldContext.StringValue() + "." + key.Name)
}
if !isIntrinsicOrCalculatedField {
intrinsicOrCalculatedField, isIntrinsicOrCalculatedField = lookupIntrinsicOrCalculatedField(key.Name)
if _, ok := tracestelemetryschema.IntrinsicFields[key.Name]; ok {
isIntrinsicOrCalculatedField = true
intrinsicOrCalculatedField = tracestelemetryschema.IntrinsicFields[key.Name]
} else if _, ok := tracestelemetryschema.CalculatedFields[key.Name]; ok {
isIntrinsicOrCalculatedField = true
intrinsicOrCalculatedField = tracestelemetryschema.CalculatedFields[key.Name]
} else if _, ok := tracestelemetryschema.IntrinsicFieldsDeprecated[key.Name]; ok {
isIntrinsicOrCalculatedField = true
intrinsicOrCalculatedField = tracestelemetryschema.IntrinsicFieldsDeprecated[key.Name]
} else if _, ok := tracestelemetryschema.CalculatedFieldsDeprecated[key.Name]; ok {
isIntrinsicOrCalculatedField = true
intrinsicOrCalculatedField = tracestelemetryschema.CalculatedFieldsDeprecated[key.Name]
}
if isIntrinsicOrCalculatedField {
@@ -292,24 +294,6 @@ func adjustTraceKey(key *telemetrytypes.TelemetryFieldKey, keys map[string][]*te
return actions
}
// lookupIntrinsicOrCalculatedField returns the intrinsic or calculated field registered under
// name, across the current and deprecated tables.
func lookupIntrinsicOrCalculatedField(name string) (telemetrytypes.TelemetryFieldKey, bool) {
if f, ok := tracestelemetryschema.IntrinsicFields[name]; ok {
return f, true
}
if f, ok := tracestelemetryschema.CalculatedFields[name]; ok {
return f, true
}
if f, ok := tracestelemetryschema.IntrinsicFieldsDeprecated[name]; ok {
return f, true
}
if f, ok := tracestelemetryschema.CalculatedFieldsDeprecated[name]; ok {
return f, true
}
return telemetrytypes.TelemetryFieldKey{}, false
}
// buildListQuery builds a query for list panel type.
func (b *traceQueryStatementBuilder) buildListQuery(
ctx context.Context,

View File

@@ -374,94 +374,6 @@ func TestStatementBuilder(t *testing.T) {
},
expectedErr: nil,
},
{
name: "scope.name filter and group by",
requestType: qbtypes.RequestTypeTimeSeries,
query: qbtypes.QueryBuilderQuery[qbtypes.TraceAggregation]{
Signal: telemetrytypes.SignalTraces,
StepInterval: qbtypes.Step{Duration: 30 * time.Second},
Aggregations: []qbtypes.TraceAggregation{
{
Expression: "count()",
},
},
Filter: &qbtypes.Filter{
Expression: "scope.name = 'opentelemetry-io'",
},
Limit: 10,
GroupBy: []qbtypes.GroupByKey{
{
TelemetryFieldKey: telemetrytypes.TelemetryFieldKey{
Name: "scope.name",
FieldContext: telemetrytypes.FieldContextScope,
},
},
},
},
expected: qbtypes.Statement{
Query: "WITH __limit_cte AS (SELECT toString(multiIf(scope.name::String <> '', scope.name::String, NULL)) AS `__GROUP_BY_KEY_0_scope.name`, count() AS __result_0 FROM signoz_traces.distributed_signoz_index_v3 WHERE (scope.name::String = ? AND scope.name::String <> '') AND timestamp >= ? AND timestamp < ? AND ts_bucket_start >= ? AND ts_bucket_start <= ? GROUP BY `__GROUP_BY_KEY_0_scope.name` ORDER BY __result_0 DESC LIMIT ?) SELECT toStartOfInterval(timestamp, INTERVAL 30 SECOND) AS ts, toString(multiIf(scope.name::String <> '', scope.name::String, NULL)) AS `__GROUP_BY_KEY_0_scope.name`, count() AS __result_0 FROM signoz_traces.distributed_signoz_index_v3 WHERE (scope.name::String = ? AND scope.name::String <> '') AND timestamp >= ? AND timestamp < ? AND ts_bucket_start >= ? AND ts_bucket_start <= ? AND (`__GROUP_BY_KEY_0_scope.name`) GLOBAL IN (SELECT `__GROUP_BY_KEY_0_scope.name` FROM __limit_cte) GROUP BY ts, `__GROUP_BY_KEY_0_scope.name`",
Args: []any{"opentelemetry-io", "1747947419000000000", "1747983448000000000", uint64(1747945619), uint64(1747983448), 10, "opentelemetry-io", "1747947419000000000", "1747983448000000000", uint64(1747945619), uint64(1747983448)},
},
expectedErr: nil,
},
{
name: "scope.version filter with scope.name group by",
requestType: qbtypes.RequestTypeTimeSeries,
query: qbtypes.QueryBuilderQuery[qbtypes.TraceAggregation]{
Signal: telemetrytypes.SignalTraces,
StepInterval: qbtypes.Step{Duration: 30 * time.Second},
Aggregations: []qbtypes.TraceAggregation{
{
Expression: "count()",
},
},
Filter: &qbtypes.Filter{
Expression: "scope.version = '1.0.0'",
},
Limit: 10,
GroupBy: []qbtypes.GroupByKey{
{
TelemetryFieldKey: telemetrytypes.TelemetryFieldKey{
Name: "scope.name",
FieldContext: telemetrytypes.FieldContextScope,
},
},
},
},
expected: qbtypes.Statement{
Query: "WITH __limit_cte AS (SELECT toString(multiIf(scope.name::String <> '', scope.name::String, NULL)) AS `__GROUP_BY_KEY_0_scope.name`, count() AS __result_0 FROM signoz_traces.distributed_signoz_index_v3 WHERE (scope.version::String = ? AND scope.version::String <> '') AND timestamp >= ? AND timestamp < ? AND ts_bucket_start >= ? AND ts_bucket_start <= ? GROUP BY `__GROUP_BY_KEY_0_scope.name` ORDER BY __result_0 DESC LIMIT ?) SELECT toStartOfInterval(timestamp, INTERVAL 30 SECOND) AS ts, toString(multiIf(scope.name::String <> '', scope.name::String, NULL)) AS `__GROUP_BY_KEY_0_scope.name`, count() AS __result_0 FROM signoz_traces.distributed_signoz_index_v3 WHERE (scope.version::String = ? AND scope.version::String <> '') AND timestamp >= ? AND timestamp < ? AND ts_bucket_start >= ? AND ts_bucket_start <= ? AND (`__GROUP_BY_KEY_0_scope.name`) GLOBAL IN (SELECT `__GROUP_BY_KEY_0_scope.name` FROM __limit_cte) GROUP BY ts, `__GROUP_BY_KEY_0_scope.name`",
Args: []any{"1.0.0", "1747947419000000000", "1747983448000000000", uint64(1747945619), uint64(1747983448), 10, "1.0.0", "1747947419000000000", "1747983448000000000", uint64(1747945619), uint64(1747983448)},
},
expectedErr: nil,
},
{
name: "scope.version filter only (no scope field in group by)",
requestType: qbtypes.RequestTypeTimeSeries,
query: qbtypes.QueryBuilderQuery[qbtypes.TraceAggregation]{
Signal: telemetrytypes.SignalTraces,
StepInterval: qbtypes.Step{Duration: 30 * time.Second},
Aggregations: []qbtypes.TraceAggregation{
{
Expression: "count()",
},
},
Filter: &qbtypes.Filter{
Expression: "scope.version = '1.0.0'",
},
Limit: 10,
GroupBy: []qbtypes.GroupByKey{
{
TelemetryFieldKey: telemetrytypes.TelemetryFieldKey{
Name: "service.name",
},
},
},
},
expected: qbtypes.Statement{
Query: "WITH __limit_cte AS (SELECT toString(multiIf(multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL, multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL), NULL)) AS `__GROUP_BY_KEY_0_service.name`, count() AS __result_0 FROM signoz_traces.distributed_signoz_index_v3 WHERE (scope.version::String = ? AND scope.version::String <> '') AND timestamp >= ? AND timestamp < ? AND ts_bucket_start >= ? AND ts_bucket_start <= ? GROUP BY `__GROUP_BY_KEY_0_service.name` ORDER BY __result_0 DESC LIMIT ?) SELECT toStartOfInterval(timestamp, INTERVAL 30 SECOND) AS ts, toString(multiIf(multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL, multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL), NULL)) AS `__GROUP_BY_KEY_0_service.name`, count() AS __result_0 FROM signoz_traces.distributed_signoz_index_v3 WHERE (scope.version::String = ? AND scope.version::String <> '') AND timestamp >= ? AND timestamp < ? AND ts_bucket_start >= ? AND ts_bucket_start <= ? AND (`__GROUP_BY_KEY_0_service.name`) GLOBAL IN (SELECT `__GROUP_BY_KEY_0_service.name` FROM __limit_cte) GROUP BY ts, `__GROUP_BY_KEY_0_service.name`",
Args: []any{"1.0.0", "1747947419000000000", "1747983448000000000", uint64(1747945619), uint64(1747983448), 10, "1.0.0", "1747947419000000000", "1747983448000000000", uint64(1747945619), uint64(1747983448)},
},
},
}
fl := flaggertest.New(t)
@@ -888,143 +800,6 @@ func TestStatementBuilderListQueryWithCorruptData(t *testing.T) {
},
expectedErr: nil,
},
{
name: "List query with scope filter only (no scope in select or group by)",
requestType: qbtypes.RequestTypeRaw,
keysMap: map[string][]*telemetrytypes.TelemetryFieldKey{
"scope.version": {
{
Name: "scope.version",
Signal: telemetrytypes.SignalTraces,
FieldContext: telemetrytypes.FieldContextScope,
FieldDataType: telemetrytypes.FieldDataTypeString,
},
},
},
query: qbtypes.QueryBuilderQuery[qbtypes.TraceAggregation]{
Signal: telemetrytypes.SignalTraces,
StepInterval: qbtypes.Step{Duration: 30 * time.Second},
Filter: &qbtypes.Filter{
Expression: "scope.version = '1.0.0'",
},
Limit: 10,
},
expected: qbtypes.Statement{
Query: "SELECT timestamp AS `__SELECT_KEY_0_timestamp`, trace_id AS `__SELECT_KEY_1_trace_id`, span_id AS `__SELECT_KEY_2_span_id`, trace_state AS `__SELECT_KEY_3_trace_state`, parent_span_id AS `__SELECT_KEY_4_parent_span_id`, flags AS `__SELECT_KEY_5_flags`, name AS `__SELECT_KEY_6_name`, kind AS `__SELECT_KEY_7_kind`, kind_string AS `__SELECT_KEY_8_kind_string`, duration_nano AS `__SELECT_KEY_9_duration_nano`, status_code AS `__SELECT_KEY_10_status_code`, status_message AS `__SELECT_KEY_11_status_message`, status_code_string AS `__SELECT_KEY_12_status_code_string`, events AS `__SELECT_KEY_13_events`, links AS `__SELECT_KEY_14_links`, response_status_code AS `__SELECT_KEY_15_response_status_code`, external_http_url AS `__SELECT_KEY_16_external_http_url`, http_url AS `__SELECT_KEY_17_http_url`, external_http_method AS `__SELECT_KEY_18_external_http_method`, http_method AS `__SELECT_KEY_19_http_method`, http_host AS `__SELECT_KEY_20_http_host`, db_name AS `__SELECT_KEY_21_db_name`, db_operation AS `__SELECT_KEY_22_db_operation`, has_error AS `__SELECT_KEY_23_has_error`, is_remote AS `__SELECT_KEY_24_is_remote`, attributes_string, attributes_number, attributes_bool, resources_string FROM signoz_traces.distributed_signoz_index_v3 WHERE (scope.version::String = ? AND scope.version::String <> '') AND timestamp >= ? AND timestamp < ? AND ts_bucket_start >= ? AND ts_bucket_start <= ? LIMIT ?",
Args: []any{"1.0.0", "1747947419000000000", "1747983448000000000", uint64(1747945619), uint64(1747983448), 10},
},
},
{
// Regression test: scope.version in selectFields with no metadata (isColumn=true filters it out)
// must still produce scope.version::String, not scope.attributes.version::String
name: "scope.version in selectFields only, no metadata (intrinsic field fallback)",
requestType: qbtypes.RequestTypeRaw,
keysMap: map[string][]*telemetrytypes.TelemetryFieldKey{},
query: qbtypes.QueryBuilderQuery[qbtypes.TraceAggregation]{
Signal: telemetrytypes.SignalTraces,
StepInterval: qbtypes.Step{Duration: 30 * time.Second},
Filter: &qbtypes.Filter{},
SelectFields: []telemetrytypes.TelemetryFieldKey{
{Name: "scope.version", FieldContext: telemetrytypes.FieldContextUnspecified},
},
Limit: 10,
},
expected: qbtypes.Statement{
Query: "SELECT timestamp AS `__SELECT_KEY_0_timestamp`, trace_id AS `__SELECT_KEY_1_trace_id`, span_id AS `__SELECT_KEY_2_span_id`, multiIf(scope.version::String <> '', scope.version::String, NULL) AS `__SELECT_KEY_3_scope.version` FROM signoz_traces.distributed_signoz_index_v3 WHERE timestamp >= ? AND timestamp < ? AND ts_bucket_start >= ? AND ts_bucket_start <= ? LIMIT ?",
Args: []any{"1747947419000000000", "1747983448000000000", uint64(1747945619), uint64(1747983448), 10},
},
},
{
// A scope attribute whose own name literally carries a `scope.` prefix (`scope.prefixed`,
// normalized to {prefixed, scope}) resolves to that attribute in a SELECT even without a
// filter: getKeySelectors emits the reconstructed `scope.prefixed` selector so the metadata
// fetch surfaces it and AdjustKey recovers the full name. Without it the `scope.` prefix is
// lost and it wrongly reads `scope.attributes.prefixed`.
name: "scope-prefixed attribute in selectFields resolves without a filter",
requestType: qbtypes.RequestTypeRaw,
keysMap: map[string][]*telemetrytypes.TelemetryFieldKey{
"scope.prefixed": {
{
Name: "scope.prefixed",
Signal: telemetrytypes.SignalTraces,
FieldContext: telemetrytypes.FieldContextScope,
FieldDataType: telemetrytypes.FieldDataTypeString,
},
},
},
query: qbtypes.QueryBuilderQuery[qbtypes.TraceAggregation]{
Signal: telemetrytypes.SignalTraces,
StepInterval: qbtypes.Step{Duration: 30 * time.Second},
Filter: &qbtypes.Filter{},
SelectFields: []telemetrytypes.TelemetryFieldKey{
{Name: "prefixed", FieldContext: telemetrytypes.FieldContextScope},
},
Limit: 10,
},
expected: qbtypes.Statement{
Query: "SELECT timestamp AS `__SELECT_KEY_0_timestamp`, trace_id AS `__SELECT_KEY_1_trace_id`, span_id AS `__SELECT_KEY_2_span_id`, multiIf(scope.attributes.`scope.prefixed` IS NOT NULL, scope.attributes.`scope.prefixed`::String, NULL) AS `__SELECT_KEY_3_scope.prefixed` FROM signoz_traces.distributed_signoz_index_v3 WHERE timestamp >= ? AND timestamp < ? AND ts_bucket_start >= ? AND ts_bucket_start <= ? LIMIT ?",
Args: []any{"1747947419000000000", "1747983448000000000", uint64(1747945619), uint64(1747983448), 10},
},
},
{
// A scope-context key whose name matches a declared scope path resolves to that
// declared path (scope.name), not the span `name` column and not an undeclared
// scope attribute. getTracesKeys surfaces the declared path as an intrinsic key
// (metadata.go), which shadows the same-named span intrinsic.
name: "scope-context name resolves to the declared scope path",
requestType: qbtypes.RequestTypeRaw,
keysMap: map[string][]*telemetrytypes.TelemetryFieldKey{
"scope.name": {
{
Name: "scope.name",
Signal: telemetrytypes.SignalTraces,
FieldContext: telemetrytypes.FieldContextScope,
FieldDataType: telemetrytypes.FieldDataTypeString,
},
},
"name": {
{
Name: "name",
Signal: telemetrytypes.SignalTraces,
FieldContext: telemetrytypes.FieldContextSpan,
FieldDataType: telemetrytypes.FieldDataTypeString,
},
},
},
query: qbtypes.QueryBuilderQuery[qbtypes.TraceAggregation]{
Signal: telemetrytypes.SignalTraces,
StepInterval: qbtypes.Step{Duration: 30 * time.Second},
Filter: &qbtypes.Filter{},
SelectFields: []telemetrytypes.TelemetryFieldKey{
{Name: "name", FieldContext: telemetrytypes.FieldContextScope},
},
Limit: 10,
},
expected: qbtypes.Statement{
Query: "SELECT timestamp AS `__SELECT_KEY_0_timestamp`, trace_id AS `__SELECT_KEY_1_trace_id`, span_id AS `__SELECT_KEY_2_span_id`, multiIf(scope.name::String <> '', scope.name::String, NULL) AS `__SELECT_KEY_3_name` FROM signoz_traces.distributed_signoz_index_v3 WHERE timestamp >= ? AND timestamp < ? AND ts_bucket_start >= ? AND ts_bucket_start <= ? LIMIT ?",
Args: []any{"1747947419000000000", "1747983448000000000", uint64(1747945619), uint64(1747983448), 10},
},
},
{
// span.scope.name (span context, name "scope.name") resolves to the declared
// scope path scope.name, not a span attribute literally named scope.name.
name: "span-context scope.name in selectFields resolves to the declared scope path",
requestType: qbtypes.RequestTypeRaw,
keysMap: map[string][]*telemetrytypes.TelemetryFieldKey{},
query: qbtypes.QueryBuilderQuery[qbtypes.TraceAggregation]{
Signal: telemetrytypes.SignalTraces,
StepInterval: qbtypes.Step{Duration: 30 * time.Second},
Filter: &qbtypes.Filter{},
SelectFields: []telemetrytypes.TelemetryFieldKey{
{Name: "scope.name", FieldContext: telemetrytypes.FieldContextSpan},
},
Limit: 10,
},
expected: qbtypes.Statement{
Query: "SELECT timestamp AS `__SELECT_KEY_0_timestamp`, trace_id AS `__SELECT_KEY_1_trace_id`, span_id AS `__SELECT_KEY_2_span_id`, multiIf(scope.name::String <> '', scope.name::String, NULL) AS `__SELECT_KEY_3_scope.name` FROM signoz_traces.distributed_signoz_index_v3 WHERE timestamp >= ? AND timestamp < ? AND ts_bucket_start >= ? AND ts_bucket_start <= ? LIMIT ?",
Args: []any{"1747947419000000000", "1747983448000000000", uint64(1747945619), uint64(1747983448), 10},
},
},
}
for _, c := range cases {

View File

@@ -180,7 +180,7 @@ func (t *telemetryMetaStore) getTracesKeys(ctx context.Context, fieldKeySelector
`CASE
// WHEN tagType = 'spanfield' THEN 1
WHEN tagType = 'resource' THEN 2
WHEN tagType = 'scope' THEN 3
// WHEN tagType = 'scope' THEN 3
WHEN tagType = 'tag' THEN 4
ELSE 5
END as priority`,

View File

@@ -4,6 +4,7 @@ import (
"context"
"fmt"
"regexp"
"strings"
schema "github.com/SigNoz/signoz-otel-collector/cmd/signozschemamigrator/schema_migrator"
"github.com/SigNoz/signoz/pkg/errors"
@@ -74,6 +75,35 @@ func (c *conditionBuilder) conditionForSearch(
return []string{sb.Or(conditions...)}, nil, nil
}
// numberAttributeIndexPredicate returns what an equality on a numeric attribute implies over
// mapValues(attributes_number), which its bloom filter indexes while the subscript the comparison
// reads matches nothing. The paired mapContains is what makes membership hold for the zero default.
func numberAttributeIndexPredicate(columns []*schema.Column, value any, sb *sqlbuilder.SelectBuilder) string {
if len(columns) != 1 || columns[0].Name != LogsV2AttributesNumberColumn {
return ""
}
// a non-numeric value means the collision handler compared the column as text, where an
// Array(Float64) membership check has no supertype
switch value.(type) {
case int, int8, int16, int32, int64, uint, uint8, uint16, uint32, uint64, float32, float64:
return fmt.Sprintf("has(mapValues(%s), %s)", LogsV2AttributesNumberColumn, sb.Var(value))
}
return ""
}
// stringAttributeIndexPredicate returns the raw-value match a case-insensitive one implies when
// the pattern holds no ASCII letter, LOWER being the identity on those bytes. A letter breaks it:
// an `a` in the pattern may have come from an `A` in the value.
func stringAttributeIndexPredicate(columns []*schema.Column, fieldExpression, pattern string, sb *sqlbuilder.SelectBuilder) string {
if len(columns) != 1 || columns[0].Name != LogsV2AttributesStringColumn {
return ""
}
if strings.ContainsFunc(pattern, func(r rune) bool { return (r >= 'a' && r <= 'z') || (r >= 'A' && r <= 'Z') }) {
return ""
}
return sb.Like(fieldExpression, pattern)
}
// isBodyJSONSearch reports whether a key addresses a path within the body JSON. Only
// an explicit Body context qualifies; a bare, context-less `body` (e.g. full-text
// `count_distinct(body)` or `body EXISTS`) is a full-text match, not a `$.body` path.
@@ -105,14 +135,14 @@ func (c *conditionBuilder) conditionForArrayFunction(
"function `%s` supports only body JSON search", operator.FunctionName()).WithUrl(functionBodyJSONSearchDocURL)
}
needle := value
element := value
if args, ok := value.([]any); ok && len(args) > 0 {
needle = args[0]
element = args[0]
}
if c.fl.BooleanOrEmpty(ctx, flagger.FeatureUseJSONBody, featuretypes.NewFlaggerEvaluationContext(orgID)) {
// JSON access plan: data-type collision handling, nested array paths.
valueType, needle := InferDataType(needle, operator, key)
valueType, element := InferDataType(element, operator, key)
// A not-found (synthesized) body path carries no metadata plan; build an exhaustive
// one so the query runs against the underlying data (with the not-found warning)
// instead of erroring, matching the regular-operator path.
@@ -123,23 +153,44 @@ func (c *conditionBuilder) conditionForArrayFunction(
}
key = keyCopy
}
return NewJSONConditionBuilder(key, valueType).buildArrayFunctionCondition(operator, needle, sb)
return NewJSONConditionBuilder(key, valueType).buildArrayFunctionCondition(operator, element, sb)
}
// legacy string-body path: type-matched array extraction, OR-ed with a scalar comparison
// for a scalar body value (coalesced to false so NOT has() matches missing-key rows).
elemType := legacyElemType(needle)
return c.legacyArrayFunctionCondition(key, operator, element, sb), nil
}
// legacyArrayFunctionCondition builds the has-family comparison over the plain body string, with
// what it implies over the indexed LOWER(body): the path, since the extraction yields nothing for an
// absent one, and each element, since it must appear in the text. The comparison decides the row.
func (c *conditionBuilder) legacyArrayFunctionCondition(
key *telemetrytypes.TelemetryFieldKey,
operator qbtypes.FilterOperator,
element any,
sb *sqlbuilder.SelectBuilder,
) string {
// type-matched array extraction, OR-ed with a scalar comparison for a scalar body value
// (coalesced to false so NOT has() matches missing-key rows).
elemType := legacyElemType(element)
arrayExpr := getBodyJSONArrayKey(key, elemType)
scalarExpr, scalarGuard, hasScalar := getBodyJSONScalarKey(key, elemType)
if list, ok := needle.([]any); ok {
vals := make([]any, len(list))
for i, v := range list {
vals[i] = legacyCoerceNeedle(v, elemType)
}
// Pin the needle array type to the haystack; scalar fallback below coerces value-level.
arrayCond := fmt.Sprintf("%s(%s, %s)", operator.FunctionName(), arrayExpr, castNeedleArray(elemType, sb.Var(vals)))
elements, isList := element.([]any)
if !isList {
elements = []any{element}
}
vals := make([]any, len(elements))
for i, v := range elements {
vals[i] = legacyCoerceElement(v, elemType)
}
var cond string
switch {
case isList:
// Pin the element array type to the array it is tested against; scalar fallback below coerces value-level.
arrayCond := fmt.Sprintf("%s(%s, %s)", operator.FunctionName(), arrayExpr, castElementArray(elemType, sb.Var(vals)))
if !hasScalar {
return arrayCond, nil
cond = arrayCond
break
}
var membership string
if operator == qbtypes.FilterOperatorHasAll {
@@ -151,19 +202,75 @@ func (c *conditionBuilder) conditionForArrayFunction(
} else {
membership = sb.In(scalarExpr, vals...)
}
return fmt.Sprintf("(%s OR ifNull(%s, false))", arrayCond, sb.And(membership, scalarGuard)), nil
cond = fmt.Sprintf("(%s OR ifNull(%s, false))", arrayCond, sb.And(membership, scalarGuard))
default:
arrayCond := fmt.Sprintf("%s(%s, %s)", operator.FunctionName(), arrayExpr, sb.Var(vals[0]))
if !hasScalar {
cond = arrayCond
break
}
cond = fmt.Sprintf("(%s OR ifNull(%s, false))", arrayCond, sb.And(sb.E(scalarExpr, vals[0]), scalarGuard))
}
typedNeedle := legacyCoerceNeedle(needle, elemType)
arrayCond := fmt.Sprintf("%s(%s, %s)", operator.FunctionName(), arrayExpr, sb.Var(typedNeedle))
if !hasScalar {
return arrayCond, nil
predicates := []string{cond}
if predicate := bodyIndexPredicate(bodyPathLiterals(key), sb); predicate != "" {
predicates = append(predicates, predicate)
}
return fmt.Sprintf("(%s OR ifNull(%s, false))", arrayCond, sb.And(sb.E(scalarExpr, typedNeedle), scalarGuard)), nil
// only strings imply text: the family compares at the element type it infers, so a quoted
// number is still a number here and its digits need not appear in the body
if elemType == telemetrytypes.FieldDataTypeString {
predicates = append(predicates, elementLiterals(operator, elements, sb)...)
}
if len(predicates) > 1 {
return sb.And(predicates...)
}
return cond
}
// castNeedleArray pins an Int64 needle array to Array(Int64) so it matches the Array(Nullable(Int64))
// haystack; without it a needle >= 2^32 binds as Array(UInt64) and hasAny/hasAll error (code 386).
func castNeedleArray(elemType telemetrytypes.FieldDataType, arg string) string {
// elementLiterals returns what the elements of a has-family filter imply about the body text. has
// and hasAll require every one, so each becomes its own predicate; hasAny requires only one, so the
// arms are ORed - and an element yielding no literal leaves that OR unassertable.
func elementLiterals(operator qbtypes.FilterOperator, elements []any, sb *sqlbuilder.SelectBuilder) []string {
if operator == qbtypes.FilterOperatorHasAny {
// resolve every element before binding anything: one unusable element voids the whole OR
runSets := make([][]string, 0, len(elements))
for _, v := range elements {
str, ok := v.(string)
if !ok {
return nil
}
runs := jsonTextRuns(str)
if len(runs) == 0 {
return nil
}
runSets = append(runSets, runs)
}
arms := make([]string, 0, len(runSets))
for _, runs := range runSets {
arms = append(arms, bodyIndexPredicate(runs, sb))
}
if len(arms) == 0 {
return nil
}
return []string{sb.Or(arms...)}
}
var predicates []string
for _, v := range elements {
str, ok := v.(string)
if !ok {
continue
}
if predicate := bodyIndexPredicate(jsonTextRuns(str), sb); predicate != "" {
predicates = append(predicates, predicate)
}
}
return predicates
}
// castElementArray pins an Int64 element array to Array(Int64) so it matches the Array(Nullable(Int64))
// it is tested against; without it an element >= 2^32 binds as Array(UInt64) and hasAny/hasAll error (code 386).
func castElementArray(elemType telemetrytypes.FieldDataType, arg string) string {
if elemType == telemetrytypes.FieldDataTypeInt64 {
return fmt.Sprintf("CAST(%s AS Array(Int64))", arg)
}
@@ -191,24 +298,24 @@ func (c *conditionBuilder) conditionForHasToken(
value any,
sb *sqlbuilder.SelectBuilder,
) (string, error) {
// hasToken takes a single needle; unwrap it from the function-argument slice.
needle := value
// hasToken takes a single token; unwrap it from the function-argument slice.
token := value
if args, ok := value.([]any); ok && len(args) > 0 {
needle = args[0]
token = args[0]
}
// hasToken matches string tokens only.
needleStr, ok := needle.(string)
tokenStr, ok := token.(string)
if !ok {
return "", errors.NewInvalidInputf(errors.CodeInvalidInput,
"function `hasToken` expects value parameter to be a string").WithUrl(hasTokenFunctionDocURL)
}
// A multi-token needle makes CH hasToken error (code 36); reject up front as a 400. Both modes flow here.
if sep, found := firstTokenSeparator(needleStr); found {
// A multi-token value makes CH hasToken error (code 36); reject up front as a 400. Both modes flow here.
if sep, found := firstTokenSeparator(tokenStr); found {
return "", errors.NewInvalidInputf(errors.CodeInvalidInput,
"function `hasToken` matches a single whole token, but %q contains the separator %q; use a substring filter (e.g. `body CONTAINS '%s'`) to search across separators",
needleStr, sep, needleStr).WithUrl(hasTokenFunctionDocURL)
tokenStr, sep, tokenStr).WithUrl(hasTokenFunctionDocURL)
}
bodyJSONEnabled := c.fl.BooleanOrEmpty(ctx, flagger.FeatureUseJSONBody, featuretypes.NewFlaggerEvaluationContext(orgID))
@@ -219,7 +326,7 @@ func (c *conditionBuilder) conditionForHasToken(
return "", errors.NewInvalidInputf(errors.CodeInvalidInput,
"function `hasToken` only supports body field as first parameter").WithUrl(hasTokenFunctionDocURL)
}
return fmt.Sprintf("hasToken(LOWER(%s), LOWER(%s))", LogsV2BodyColumn, sb.Var(needle)), nil
return fmt.Sprintf("hasToken(LOWER(%s), LOWER(%s))", LogsV2BodyColumn, sb.Var(token)), nil
}
// JSON mode: a bare body/body.message key searches the body.message column; any other body
@@ -228,7 +335,7 @@ func (c *conditionBuilder) conditionForHasToken(
// falls through and emits dynamicElement over the already-typed String column, which errors.
if key.Name == LogsV2BodyColumn || key.Name == bodyMessageField ||
(key.FieldContext == telemetrytypes.FieldContextBody && key.Name == messageSubField) {
return fmt.Sprintf("hasToken(LOWER(%s), LOWER(%s))", bodyMessageField, sb.Var(needle)), nil
return fmt.Sprintf("hasToken(LOWER(%s), LOWER(%s))", bodyMessageField, sb.Var(token)), nil
}
if key.FieldContext == telemetrytypes.FieldContextBody {
// A not-found (synthesized) body path carries no metadata plan; build an exhaustive
@@ -240,7 +347,7 @@ func (c *conditionBuilder) conditionForHasToken(
}
key = keyCopy
}
return NewJSONConditionBuilder(key, telemetrytypes.FieldDataTypeString).buildTokenFunctionCondition(needle, sb)
return NewJSONConditionBuilder(key, telemetrytypes.FieldDataTypeString).buildTokenFunctionCondition(token, sb)
}
return "", errors.NewInvalidInputf(errors.CodeInvalidInput,
"function `hasToken` only supports the body field or a body JSON string field as first parameter").WithUrl(hasTokenFunctionDocURL)
@@ -269,6 +376,9 @@ func (c *conditionBuilder) conditionForResolvedKey(
return "", err
}
useJSONBody := c.fl.BooleanOrEmpty(ctx, flagger.FeatureUseJSONBody, featuretypes.NewFlaggerEvaluationContext(orgID))
legacyBodyJSONSearch := isBodyJSONSearch(key, columns) && !useJSONBody
// has/hasAny/hasAll take the body-JSON path, not the normal operator paths.
if operator.IsArrayFunctionOperator() {
return c.conditionForArrayFunction(ctx, orgID, key, operator, value, columns, sb)
@@ -276,7 +386,7 @@ func (c *conditionBuilder) conditionForResolvedKey(
// TODO(Piyush): Update this to support multiple JSON columns based on evolutions
for _, column := range columns {
if column.Type.GetType() == schema.ColumnTypeEnumJSON && isBodyJSONSearch(key, columns) && c.fl.BooleanOrEmpty(ctx, flagger.FeatureUseJSONBody, featuretypes.NewFlaggerEvaluationContext(orgID)) && key.Name != messageSubField {
if column.Type.GetType() == schema.ColumnTypeEnumJSON && isBodyJSONSearch(key, columns) && useJSONBody && key.Name != messageSubField {
valueType, value := InferDataType(value, operator, key)
if len(key.JSONPlan) == 0 {
keyCopy := telemetrytypes.NewTelemetryFieldKey(key.Name, key.FieldContext, key.FieldDataType)
@@ -305,19 +415,83 @@ func (c *conditionBuilder) conditionForResolvedKey(
}
// Check if this is a body JSON search (legacy string-body path, JSON flag off).
if isBodyJSONSearch(key, columns) && !c.fl.BooleanOrEmpty(ctx, flagger.FeatureUseJSONBody, featuretypes.NewFlaggerEvaluationContext(orgID)) {
fieldExpression, value = GetBodyJSONKey(ctx, key, operator, value)
if legacyBodyJSONSearch {
return c.conditionForLegacyBodyJSON(ctx, orgID, startNs, endNs, key, operator, value, columns, sb)
}
fieldExpression, value = querybuilder.DataTypeCollisionHandledFieldName(key, value, fieldExpression, operator)
return c.conditionForOperator(ctx, orgID, startNs, endNs, key, operator, value, columns, fieldExpression, sb)
}
// conditionForLegacyBodyJSON renders a filter over a path inside the plain string body, with what it
// implies over the indexed LOWER(body) - nothing such a filter compares matches an index expression.
func (c *conditionBuilder) conditionForLegacyBodyJSON(
ctx context.Context,
orgID valuer.UUID,
startNs, endNs uint64,
key *telemetrytypes.TelemetryFieldKey,
operator qbtypes.FilterOperator,
value any,
columns []*schema.Column,
sb *sqlbuilder.SelectBuilder,
) (string, error) {
if operator == qbtypes.FilterOperatorExists || operator == qbtypes.FilterOperatorNotExists {
exists := GetBodyJSONKeyForExists(ctx, key, operator, value)
if operator == qbtypes.FilterOperatorNotExists {
// matches the rows without the path, which say nothing about the body text
return "NOT " + exists, nil
}
if predicate := bodyIndexPredicate(bodyPathLiterals(key), sb); predicate != "" {
return sb.And(exists, predicate), nil
}
return exists, nil
}
fieldExpression, value := GetBodyJSONKey(ctx, key, operator, value)
fieldExpression, value = querybuilder.DataTypeCollisionHandledFieldName(key, value, fieldExpression, operator)
cond, err := c.conditionForOperator(ctx, orgID, startNs, endNs, key, operator, value, columns, fieldExpression, sb)
if err != nil {
return "", err
}
if predicate := bodyIndexPredicate(bodyValueLiterals(operator, value), sb); predicate != "" {
return sb.And(cond, predicate), nil
}
return cond, nil
}
// conditionForOperator renders the comparison itself, once the field expression and value have been
// resolved for the column the key landed on.
func (c *conditionBuilder) conditionForOperator(
ctx context.Context,
orgID valuer.UUID,
startNs, endNs uint64,
key *telemetrytypes.TelemetryFieldKey,
operator qbtypes.FilterOperator,
value any,
columns []*schema.Column,
fieldExpression string,
sb *sqlbuilder.SelectBuilder,
) (string, error) {
// make use of case insensitive index for body
if fieldExpression == "body" || fieldExpression == messageSubColumn {
switch operator {
case qbtypes.FilterOperatorEqual:
// Bloom filters index lower(body), not the column; `=` still decides the row.
if _, ok := value.(string); ok && fieldExpression == LogsV2BodyColumn {
return sb.And(
sb.E(fieldExpression, value),
fmt.Sprintf("LOWER(%s) = LOWER(%s)", fieldExpression, sb.Var(value)),
), nil
}
case qbtypes.FilterOperatorLike:
return sb.ILike(fieldExpression, value), nil
case qbtypes.FilterOperatorNotLike:
return sb.NotILike(fieldExpression, value), nil
if _, ok := value.(string); ok && fieldExpression == LogsV2BodyColumn {
return sb.And(
sb.Like(fieldExpression, value),
sb.ILike(fieldExpression, value),
), nil
}
case qbtypes.FilterOperatorRegexp:
// Note: Escape $$ to $$$$ to avoid sqlbuilder interpreting materialized $ signs
// Only needed because we are using sprintf instead of sb.Match (not implemented in sqlbuilder)
@@ -333,6 +507,9 @@ func (c *conditionBuilder) conditionForResolvedKey(
switch operator {
// regular operators
case qbtypes.FilterOperatorEqual:
if predicate := numberAttributeIndexPredicate(columns, value, sb); predicate != "" {
return sb.And(sb.E(fieldExpression, value), predicate), nil
}
return sb.E(fieldExpression, value), nil
case qbtypes.FilterOperatorNotEqual:
return sb.NE(fieldExpression, value), nil
@@ -351,17 +528,16 @@ func (c *conditionBuilder) conditionForResolvedKey(
case qbtypes.FilterOperatorNotLike:
return sb.NotLike(fieldExpression, value), nil
case qbtypes.FilterOperatorILike:
if pattern, ok := value.(string); ok {
if predicate := stringAttributeIndexPredicate(columns, fieldExpression, pattern, sb); predicate != "" {
return sb.And(sb.ILike(fieldExpression, pattern), predicate), nil
}
}
return sb.ILike(fieldExpression, value), nil
case qbtypes.FilterOperatorNotILike:
return sb.NotILike(fieldExpression, value), nil
case qbtypes.FilterOperatorExists, qbtypes.FilterOperatorNotExists:
if isBodyJSONSearch(key, columns) && !c.fl.BooleanOrEmpty(ctx, flagger.FeatureUseJSONBody, featuretypes.NewFlaggerEvaluationContext(orgID)) {
if operator == qbtypes.FilterOperatorExists {
return GetBodyJSONKeyForExists(ctx, key, operator, value), nil
}
return "NOT " + GetBodyJSONKeyForExists(ctx, key, operator, value), nil
}
pred, err := querybuilder.ExistsExpression(columns, key, startNs, endNs, fieldExpression, operator == qbtypes.FilterOperatorExists)
if err != nil {
return "", err
@@ -369,7 +545,13 @@ func (c *conditionBuilder) conditionForResolvedKey(
return sqlbuilder.Escape(pred), nil
case qbtypes.FilterOperatorContains:
return sb.ILike(fieldExpression, fmt.Sprintf("%%%s%%", value)), nil
// The map value indexes are over raw mapValues, which a case-insensitive match reaches only
// for the patterns stringAttributeIndexPredicate can assert the raw value from.
pattern := fmt.Sprintf("%%%s%%", value)
if predicate := stringAttributeIndexPredicate(columns, fieldExpression, pattern, sb); predicate != "" {
return sb.And(sb.ILike(fieldExpression, pattern), predicate), nil
}
return sb.ILike(fieldExpression, pattern), nil
case qbtypes.FilterOperatorNotContains:
return sb.NotILike(fieldExpression, fmt.Sprintf("%%%s%%", value)), nil

View File

@@ -2,6 +2,7 @@ package logstelemetryschema
import (
"context"
schema "github.com/SigNoz/signoz-otel-collector/cmd/signozschemamigrator/schema_migrator"
"testing"
"time"
@@ -168,9 +169,9 @@ func TestConditionFor(t *testing.T) {
FieldContext: telemetrytypes.FieldContextLog,
},
operator: qbtypes.FilterOperatorEqual,
value: "error message",
expectedSQL: "body = ?",
expectedArgs: []any{"error message"},
value: "Error Message",
expectedSQL: "(body = ? AND LOWER(body) = LOWER(?))",
expectedArgs: []any{"Error Message", "Error Message"},
expectedError: nil,
},
{
@@ -207,8 +208,8 @@ func TestConditionFor(t *testing.T) {
},
operator: qbtypes.FilterOperatorLike,
value: "%error%",
expectedSQL: "LOWER(body) LIKE LOWER(?)",
expectedArgs: []any{"%error%"},
expectedSQL: "(body LIKE ? AND LOWER(body) LIKE LOWER(?))",
expectedArgs: []any{"%error%", "%error%"},
expectedError: nil,
},
{
@@ -219,7 +220,7 @@ func TestConditionFor(t *testing.T) {
},
operator: qbtypes.FilterOperatorNotLike,
value: "%error%",
expectedSQL: "LOWER(body) NOT LIKE LOWER(?)",
expectedSQL: "body NOT LIKE ?",
expectedArgs: []any{"%error%"},
expectedError: nil,
},
@@ -258,8 +259,8 @@ func TestConditionFor(t *testing.T) {
},
operator: qbtypes.FilterOperatorContains,
value: 521509198310,
expectedSQL: "LOWER(attributes_string['user.id']) LIKE LOWER(?)",
expectedArgs: []any{"%521509198310%"},
expectedSQL: "(LOWER(attributes_string['user.id']) LIKE LOWER(?) AND attributes_string['user.id'] LIKE ?)",
expectedArgs: []any{"%521509198310%", "%521509198310%"},
expectedError: nil,
},
{
@@ -619,8 +620,8 @@ func TestConditionForMultipleKeys(t *testing.T) {
},
operator: qbtypes.FilterOperatorEqual,
value: "error message",
expectedSQL: "body = ? AND severity_text = ?",
expectedArgs: []any{"error message", "error message"},
expectedSQL: "(body = ? AND LOWER(body) = LOWER(?)) AND severity_text = ?",
expectedArgs: []any{"error message", "error message", "error message"},
expectedError: nil,
},
}
@@ -906,8 +907,8 @@ func TestConditionForJSONBodySearch(t *testing.T) {
}
}
// IN on the body column routes each value back through the `=` path; the SQL it produces
// must stay what the shared IN handling produced before, including for a mixed-type list.
// IN on the body column routes each value back through the `=` path, so every arm picks up
// the lower(body) companion — including the values a mixed-type list stringifies.
func TestConditionForBodyIn(t *testing.T) {
testCases := []struct {
name string
@@ -918,14 +919,14 @@ func TestConditionForBodyIn(t *testing.T) {
{
name: "strings",
values: []any{"alpha", "beta"},
expectedSQL: "(body = ? OR body = ?)",
expectedArgs: []any{"alpha", "beta"},
expectedSQL: "((body = ? AND LOWER(body) = LOWER(?)) OR (body = ? AND LOWER(body) = LOWER(?)))",
expectedArgs: []any{"alpha", "alpha", "beta", "beta"},
},
{
name: "mixed types are stringified before they reach the column",
values: []any{"alpha", float64(1), true},
expectedSQL: "(body = ? OR body = ? OR body = ?)",
expectedArgs: []any{"alpha", "1", "true"},
expectedSQL: "((body = ? AND LOWER(body) = LOWER(?)) OR (body = ? AND LOWER(body) = LOWER(?)) OR (body = ? AND LOWER(body) = LOWER(?)))",
expectedArgs: []any{"alpha", "alpha", "1", "1", "true", "true"},
},
}
@@ -954,3 +955,228 @@ func TestConditionForBodyIn(t *testing.T) {
})
}
}
// ClickHouse treats `\` as an escape only before `%`, `_` and itself.
func TestLikePatternLiterals(t *testing.T) {
testCases := []struct {
name string
pattern string
expected []string
}{
{"contains wraps a plain value", "%error%", []string{"error"}},
{"wildcards split runs", "%foo%bar%", []string{"foo", "bar"}},
{"underscore splits too", "a_b", []string{"a", "b"}},
{"escaped wildcards stay literal", `%100\%\_off%`, []string{`100%_off`}},
{"escaped backslash collapses", `%C:\\tmp%`, []string{`C:\tmp`}},
{"backslash before other chars is literal", `%C:\tmp%`, []string{`C:\tmp`}},
{"trailing backslash is literal", `%path\`, []string{`path\`}},
{"no literals at all", "%_%", nil},
}
for _, tc := range testCases {
t.Run(tc.name, func(t *testing.T) {
assert.Equal(t, tc.expected, likePatternLiterals(tc.pattern))
})
}
}
// The literals have to hold whichever encoder wrote the body, so the runs stop at every byte
// one of them may rewrite.
func TestJSONTextRuns(t *testing.T) {
testCases := []struct {
name string
value string
expected []string
}{
{"plain text is one run", "checkout failed", []string{"checkout failed"}},
{"a run with nothing to split on is kept whole", "abc", []string{"abc"}},
{"quote splits the run", `say "hello there"`, []string{"say ", "hello there"}},
{"slash splits the run, PHP escapes it", "/api/v1/users", []string{"api", "v1", "users"}},
{"ampersand and angles split, Go escapes them", "a&b<c>dddd", []string{"a", "b", "c", "dddd"}},
{"non-ascii splits, Python escapes it", "order café latte", []string{"order caf", " latte"}},
{"newline splits", "line one\nline two", []string{"line one", "line two"}},
}
for _, tc := range testCases {
t.Run(tc.name, func(t *testing.T) {
assert.Equal(t, tc.expected, jsonTextRuns(tc.value))
})
}
}
func TestBodyPathLiterals(t *testing.T) {
testCases := []struct {
name string
key string
expected []string
}{
{"quoting lifts a short name over the ngram size", "id", []string{`"id"`}},
{"one literal per component", "response.status_code", []string{`"response"`, `"status_code"`}},
{"array suffixes are trimmed", "items[*].sku", []string{`"items"`, `"sku"`}},
{"every component is carried", "a.b.count", []string{`"a"`, `"b"`, `"count"`}},
{"a component an encoder may rewrite is dropped", "user/name.email", []string{`"email"`}},
{"nothing usable", "us/er.na/me", nil},
}
for _, tc := range testCases {
t.Run(tc.name, func(t *testing.T) {
key := telemetrytypes.NewTelemetryFieldKey(tc.key, telemetrytypes.FieldContextBody, telemetrytypes.FieldDataTypeUnspecified)
assert.Equal(t, tc.expected, bodyPathLiterals(key))
})
}
}
// The path literals ride on the existence assertion and the value literals on the comparison, so
// a filter carries each at most once. Nothing rides on a negated operator: it matches rows
// without the path, which say nothing about the body text.
func TestLegacyBodyIndexPredicates(t *testing.T) {
testCases := []struct {
name string
key string
operator qbtypes.FilterOperator
value any
expected string
expectedArgs []any
}{
{
name: "exists carries the path",
key: "user_id",
operator: qbtypes.FilterOperatorExists,
expected: `LOWER(body) LIKE LOWER(?)`,
expectedArgs: []any{`%"user\_id"%`},
},
{
name: "equality carries the value",
key: "status",
operator: qbtypes.FilterOperatorEqual,
value: "timeout_error",
expected: `LOWER(body) LIKE LOWER(?)`,
expectedArgs: []any{`%timeout\_error%`},
},
{
name: "contains carries the value",
key: "message",
operator: qbtypes.FilterOperatorContains,
value: "upstream refused",
expected: `LOWER(body) LIKE LOWER(?)`,
expectedArgs: []any{`%upstream refused%`},
},
{
name: "like carries one literal per run of the pattern",
key: "message",
operator: qbtypes.FilterOperatorLike,
value: "%conn%refused%",
expected: `LOWER(body) LIKE LOWER(?)`,
expectedArgs: []any{`%conn%refused%`},
},
{
name: "has carries the path and the element",
key: "tags[*]",
operator: qbtypes.FilterOperatorHas,
value: []any{"production"},
// The element rides on its own predicate rather than being pinned next to the key:
// has() over the extracted array says nothing about where in the text it sits.
expected: `LOWER(body) LIKE LOWER(?) AND LOWER(body) LIKE LOWER(?)`,
expectedArgs: []any{`%"tags"%`, "%production%"},
},
{
name: "hasAll carries one literal per element",
key: "tags[*]",
operator: qbtypes.FilterOperatorHasAll,
value: []any{[]any{"production", "webserver"}},
expected: `LOWER(body) LIKE LOWER(?) AND LOWER(body) LIKE LOWER(?) AND LOWER(body) LIKE LOWER(?)`,
expectedArgs: []any{`%"tags"%`, "%production%", "%webserver%"},
},
{
// hasAny asks for one of the elements, so the arms are ORed — ANDing them would
// demand every element be present.
name: "hasAny ORs the element literals",
key: "tags[*]",
operator: qbtypes.FilterOperatorHasAny,
value: []any{[]any{"production", "webserver"}},
expected: `LOWER(body) LIKE LOWER(?) AND (LOWER(body) LIKE LOWER(?) OR LOWER(body) LIKE LOWER(?))`,
expectedArgs: []any{`%"tags"%`, "%production%", "%webserver%"},
},
{
// one element with no usable literal voids the whole OR: the filter can still match
// through that element, so nothing about the text is implied.
name: "hasAny drops the OR when an element carries no literal",
key: "tags[*]",
operator: qbtypes.FilterOperatorHasAny,
value: []any{[]any{"production", "/"}},
expected: `LOWER(body) LIKE LOWER(?)`,
expectedArgs: []any{`%"tags"%`},
},
{
name: "numeric elements carry nothing",
key: "ids[*]",
operator: qbtypes.FilterOperatorHasAny,
value: []any{[]any{"9007199254740993", "9007199254740994"}},
expected: `LOWER(body) LIKE LOWER(?)`,
expectedArgs: []any{`%"ids"%`},
},
{
name: "a number carries nothing",
key: "user_id",
operator: qbtypes.FilterOperatorEqual,
value: int64(123),
},
{
// IN delegates to `=`, so the literal rides each arm rather than the IN itself
name: "IN carries one literal per arm it delegates to",
key: "status",
operator: qbtypes.FilterOperatorIn,
value: []any{"timeout_error", "conn_refused"},
expected: `(JSON_VALUE(body, '$."status"') = ? AND LOWER(body) LIKE LOWER(?)) OR (JSON_VALUE(body, '$."status"') = ? AND LOWER(body) LIKE LOWER(?))`,
expectedArgs: []any{`%timeout\_error%`, `%conn\_refused%`},
},
{
name: "not equal carries nothing",
key: "status",
operator: qbtypes.FilterOperatorNotEqual,
value: "timeout_error",
},
{
name: "not exists carries nothing",
key: "user_id",
operator: qbtypes.FilterOperatorNotExists,
},
{
name: "not contains carries nothing",
key: "message",
operator: qbtypes.FilterOperatorNotContains,
value: "upstream refused",
},
}
fl := flaggertest.New(t)
cb := NewConditionBuilder(NewFieldMapper(fl), fl)
ctx := context.Background()
columns := []*schema.Column{logsV2Columns[LogsV2BodyColumn]}
for _, tc := range testCases {
t.Run(tc.name, func(t *testing.T) {
sb := sqlbuilder.NewSelectBuilder()
sb.Select("1").From("t")
key := telemetrytypes.NewTelemetryFieldKey(tc.key, telemetrytypes.FieldContextBody, telemetrytypes.FieldDataTypeUnspecified)
var cond string
var err error
if tc.operator.IsArrayFunctionOperator() {
cond, err = cb.conditionForArrayFunction(ctx, valuer.UUID{}, key, tc.operator, tc.value, columns, sb)
} else {
cond, err = cb.conditionForLegacyBodyJSON(ctx, valuer.UUID{}, 0, 0, key, tc.operator, tc.value, columns, sb)
}
require.NoError(t, err)
sb.Where(cond)
query, args := sb.BuildWithFlavor(sqlbuilder.ClickHouse)
if tc.expected == "" {
assert.NotContains(t, query, "LOWER(body) LIKE")
return
}
assert.Contains(t, query, tc.expected)
assert.Subset(t, args, tc.expectedArgs)
})
}
}

View File

@@ -44,168 +44,169 @@ func TestFilterExprLogsBodyJSON(t *testing.T) {
category: "json",
query: "has(body.requestor_list[*], 'index_service')",
shouldPass: true,
expectedQuery: `WHERE (has(JSONExtract(JSON_QUERY(body, '$."requestor_list"[*]'), 'Array(Nullable(String))'), ?) OR ifNull((JSON_VALUE(body, '$."requestor_list"') = ? AND JSONType(body, 'requestor_list') NOT IN ('Array', 'Object', 'Null')), false))`,
expectedArgs: []any{"index_service", "index_service"},
expectedQuery: `WHERE ((has(JSONExtract(JSON_QUERY(body, '$."requestor_list"[*]'), 'Array(Nullable(String))'), ?) OR ifNull((JSON_VALUE(body, '$."requestor_list"') = ? AND JSONType(body, 'requestor_list') NOT IN ('Array', 'Object', 'Null')), false)) AND LOWER(body) LIKE LOWER(?) AND LOWER(body) LIKE LOWER(?))`,
expectedArgs: []any{"index_service", "index_service", "%\"requestor\\_list\"%", "%index\\_service%"},
expectedErrorContains: "",
},
{
category: "json",
query: "has(body.int_numbers[*], 2)",
shouldPass: true,
expectedQuery: `WHERE (has(JSONExtract(JSON_QUERY(body, '$."int_numbers"[*]'), 'Array(Nullable(Float64))'), ?) OR ifNull((JSONExtract(JSON_VALUE(body, '$."int_numbers"'), 'Nullable(Float64)') = ? AND JSONType(body, 'int_numbers') NOT IN ('Array', 'Object', 'Null')), false))`,
expectedArgs: []any{float64(2), float64(2)},
expectedQuery: `WHERE ((has(JSONExtract(JSON_QUERY(body, '$."int_numbers"[*]'), 'Array(Nullable(Float64))'), ?) OR ifNull((JSONExtract(JSON_VALUE(body, '$."int_numbers"'), 'Nullable(Float64)') = ? AND JSONType(body, 'int_numbers') NOT IN ('Array', 'Object', 'Null')), false)) AND LOWER(body) LIKE LOWER(?))`,
expectedArgs: []any{float64(2), float64(2), "%\"int\\_numbers\"%"},
expectedErrorContains: "",
},
{
category: "json",
query: "has(body.bool[*], true)",
shouldPass: true,
expectedQuery: `WHERE (has(JSONExtract(JSON_QUERY(body, '$."bool"[*]'), 'Array(Nullable(String))'), ?) OR ifNull((JSON_VALUE(body, '$."bool"') = ? AND JSONType(body, 'bool') NOT IN ('Array', 'Object', 'Null')), false))`,
expectedArgs: []any{"true", "true"},
expectedQuery: `WHERE ((has(JSONExtract(JSON_QUERY(body, '$."bool"[*]'), 'Array(Nullable(String))'), ?) OR ifNull((JSON_VALUE(body, '$."bool"') = ? AND JSONType(body, 'bool') NOT IN ('Array', 'Object', 'Null')), false)) AND LOWER(body) LIKE LOWER(?))`,
expectedArgs: []any{"true", "true", "%\"bool\"%"},
expectedErrorContains: "",
},
{
category: "json",
query: "NOT has(body.nested_num[*].float_nums[*], 2.2)",
shouldPass: true,
expectedQuery: `WHERE NOT (has(JSONExtract(JSON_QUERY(body, '$."nested_num"[*]."float_nums"[*]'), 'Array(Nullable(Float64))'), ?))`,
expectedArgs: []any{float64(2.2)},
expectedQuery: `WHERE NOT ((has(JSONExtract(JSON_QUERY(body, '$."nested_num"[*]."float_nums"[*]'), 'Array(Nullable(Float64))'), ?) AND LOWER(body) LIKE LOWER(?)))`,
expectedArgs: []any{float64(2.2), "%\"nested\\_num\"%\"float\\_nums\"%"},
expectedErrorContains: "",
},
{
category: "json",
query: "has(body.tags, 'production')",
shouldPass: true,
expectedQuery: `WHERE (has(JSONExtract(JSON_QUERY(body, '$."tags"[*]'), 'Array(Nullable(String))'), ?) OR ifNull((JSON_VALUE(body, '$."tags"') = ? AND JSONType(body, 'tags') NOT IN ('Array', 'Object', 'Null')), false))`,
expectedArgs: []any{"production", "production"},
expectedQuery: `WHERE ((has(JSONExtract(JSON_QUERY(body, '$."tags"[*]'), 'Array(Nullable(String))'), ?) OR ifNull((JSON_VALUE(body, '$."tags"') = ? AND JSONType(body, 'tags') NOT IN ('Array', 'Object', 'Null')), false)) AND LOWER(body) LIKE LOWER(?) AND LOWER(body) LIKE LOWER(?))`,
expectedArgs: []any{"production", "production", "%\"tags\"%", "%production%"},
expectedErrorContains: "",
},
{
category: "json",
query: "hasAny(body.tags, ['critical', 'test'])",
shouldPass: true,
expectedQuery: `WHERE (hasAny(JSONExtract(JSON_QUERY(body, '$."tags"[*]'), 'Array(Nullable(String))'), ?) OR ifNull((JSON_VALUE(body, '$."tags"') IN (?, ?) AND JSONType(body, 'tags') NOT IN ('Array', 'Object', 'Null')), false))`,
expectedArgs: []any{[]any{"critical", "test"}, "critical", "test"},
expectedQuery: `WHERE ((hasAny(JSONExtract(JSON_QUERY(body, '$."tags"[*]'), 'Array(Nullable(String))'), ?) OR ifNull((JSON_VALUE(body, '$."tags"') IN (?, ?) AND JSONType(body, 'tags') NOT IN ('Array', 'Object', 'Null')), false)) AND LOWER(body) LIKE LOWER(?) AND (LOWER(body) LIKE LOWER(?) OR LOWER(body) LIKE LOWER(?)))`,
expectedArgs: []any{[]any{"critical", "test"}, "critical", "test", "%\"tags\"%", "%critical%", "%test%"},
expectedErrorContains: "",
},
{
category: "json",
query: "hasAll(body.tags, ['production', 'web'])",
shouldPass: true,
expectedQuery: `WHERE (hasAll(JSONExtract(JSON_QUERY(body, '$."tags"[*]'), 'Array(Nullable(String))'), ?) OR ifNull(((JSON_VALUE(body, '$."tags"') = ? AND JSON_VALUE(body, '$."tags"') = ?) AND JSONType(body, 'tags') NOT IN ('Array', 'Object', 'Null')), false))`,
expectedArgs: []any{[]any{"production", "web"}, "production", "web"},
expectedQuery: `WHERE ((hasAll(JSONExtract(JSON_QUERY(body, '$."tags"[*]'), 'Array(Nullable(String))'), ?) OR ifNull(((JSON_VALUE(body, '$."tags"') = ? AND JSON_VALUE(body, '$."tags"') = ?) AND JSONType(body, 'tags') NOT IN ('Array', 'Object', 'Null')), false)) AND LOWER(body) LIKE LOWER(?) AND LOWER(body) LIKE LOWER(?) AND LOWER(body) LIKE LOWER(?))`,
expectedArgs: []any{[]any{"production", "web"}, "production", "web", "%\"tags\"%", "%production%", "%web%"},
expectedErrorContains: "",
},
{
category: "json",
query: "has(body.ids, \"200\")",
shouldPass: true,
expectedQuery: `WHERE (has(JSONExtract(JSON_QUERY(body, '$."ids"[*]'), 'Array(Nullable(Int64))'), ?) OR ifNull((JSONExtract(JSON_VALUE(body, '$."ids"'), 'Nullable(Int64)') = ? AND JSONType(body, 'ids') NOT IN ('Array', 'Object', 'Null')), false))`,
expectedArgs: []any{int64(200), int64(200)},
expectedQuery: `WHERE ((has(JSONExtract(JSON_QUERY(body, '$."ids"[*]'), 'Array(Nullable(Int64))'), ?) OR ifNull((JSONExtract(JSON_VALUE(body, '$."ids"'), 'Nullable(Int64)') = ? AND JSONType(body, 'ids') NOT IN ('Array', 'Object', 'Null')), false)) AND LOWER(body) LIKE LOWER(?))`,
expectedArgs: []any{int64(200), int64(200), "%\"ids\"%"},
expectedErrorContains: "",
},
{
// Big-int needle CAST to Array(Int64) to match the haystack (else 386).
// Big-int element CAST to Array(Int64) to match the array it is tested against (else 386).
category: "json",
query: `hasAny(body.ids, ['9007199254740993', '9007199254740994'])`,
shouldPass: true,
expectedQuery: `WHERE (hasAny(JSONExtract(JSON_QUERY(body, '$."ids"[*]'), 'Array(Nullable(Int64))'), CAST(? AS Array(Int64))) OR ifNull((JSONExtract(JSON_VALUE(body, '$."ids"'), 'Nullable(Int64)') IN (?, ?) AND JSONType(body, 'ids') NOT IN ('Array', 'Object', 'Null')), false))`,
expectedArgs: []any{[]any{int64(9007199254740993), int64(9007199254740994)}, int64(9007199254740993), int64(9007199254740994)},
expectedQuery: `WHERE ((hasAny(JSONExtract(JSON_QUERY(body, '$."ids"[*]'), 'Array(Nullable(Int64))'), CAST(? AS Array(Int64))) OR ifNull((JSONExtract(JSON_VALUE(body, '$."ids"'), 'Nullable(Int64)') IN (?, ?) AND JSONType(body, 'ids') NOT IN ('Array', 'Object', 'Null')), false)) AND LOWER(body) LIKE LOWER(?))`,
expectedArgs: []any{[]any{int64(9007199254740993), int64(9007199254740994)}, int64(9007199254740993), int64(9007199254740994), "%\"ids\"%"},
expectedErrorContains: "",
},
{
category: "json",
query: `hasAll(body.ids, ['9007199254740993', '9007199254740994'])`,
shouldPass: true,
expectedQuery: `WHERE (hasAll(JSONExtract(JSON_QUERY(body, '$."ids"[*]'), 'Array(Nullable(Int64))'), CAST(? AS Array(Int64))) OR ifNull(((JSONExtract(JSON_VALUE(body, '$."ids"'), 'Nullable(Int64)') = ? AND JSONExtract(JSON_VALUE(body, '$."ids"'), 'Nullable(Int64)') = ?) AND JSONType(body, 'ids') NOT IN ('Array', 'Object', 'Null')), false))`,
expectedArgs: []any{[]any{int64(9007199254740993), int64(9007199254740994)}, int64(9007199254740993), int64(9007199254740994)},
expectedQuery: `WHERE ((hasAll(JSONExtract(JSON_QUERY(body, '$."ids"[*]'), 'Array(Nullable(Int64))'), CAST(? AS Array(Int64))) OR ifNull(((JSONExtract(JSON_VALUE(body, '$."ids"'), 'Nullable(Int64)') = ? AND JSONExtract(JSON_VALUE(body, '$."ids"'), 'Nullable(Int64)') = ?) AND JSONType(body, 'ids') NOT IN ('Array', 'Object', 'Null')), false)) AND LOWER(body) LIKE LOWER(?))`,
expectedArgs: []any{[]any{int64(9007199254740993), int64(9007199254740994)}, int64(9007199254740993), int64(9007199254740994), "%\"ids\"%"},
expectedErrorContains: "",
},
{
category: "json",
query: "body.message = hello",
shouldPass: true,
expectedQuery: `WHERE (JSON_VALUE(body, '$."message"') = ? AND JSON_EXISTS(body, '$."message"'))`,
expectedArgs: []any{"hello"},
expectedQuery: `WHERE ((JSON_VALUE(body, '$."message"') = ? AND LOWER(body) LIKE LOWER(?)) AND (JSON_EXISTS(body, '$."message"') AND LOWER(body) LIKE LOWER(?)))`,
expectedArgs: []any{"hello", "%hello%", "%\"message\"%"},
expectedErrorContains: "",
},
{
category: "json",
query: "body.status = 1",
shouldPass: true,
expectedQuery: `WHERE (JSONExtract(JSON_VALUE(body, '$."status"'), 'Float64') = ? AND JSON_EXISTS(body, '$."status"'))`,
expectedArgs: []any{float64(1)},
expectedQuery: `WHERE (JSONExtract(JSON_VALUE(body, '$."status"'), 'Float64') = ? AND (JSON_EXISTS(body, '$."status"') AND LOWER(body) LIKE LOWER(?)))`,
expectedArgs: []any{float64(1), "%\"status\"%"},
expectedErrorContains: "",
},
{
category: "json",
query: "body.status = 1.1",
shouldPass: true,
expectedQuery: `WHERE (JSONExtract(JSON_VALUE(body, '$."status"'), 'Float64') = ? AND JSON_EXISTS(body, '$."status"'))`,
expectedArgs: []any{float64(1.1)},
expectedQuery: `WHERE (JSONExtract(JSON_VALUE(body, '$."status"'), 'Float64') = ? AND (JSON_EXISTS(body, '$."status"') AND LOWER(body) LIKE LOWER(?)))`,
expectedArgs: []any{float64(1.1), "%\"status\"%"},
expectedErrorContains: "",
},
{
category: "json",
query: "body.boolkey = true",
shouldPass: true,
expectedQuery: `WHERE (JSONExtract(JSON_VALUE(body, '$."boolkey"'), 'Bool') = ? AND JSON_EXISTS(body, '$."boolkey"'))`,
expectedArgs: []any{true},
expectedQuery: `WHERE (JSONExtract(JSON_VALUE(body, '$."boolkey"'), 'Bool') = ? AND (JSON_EXISTS(body, '$."boolkey"') AND LOWER(body) LIKE LOWER(?)))`,
expectedArgs: []any{true, "%\"boolkey\"%"},
expectedErrorContains: "",
},
{
category: "json",
query: "body.status > 200",
shouldPass: true,
expectedQuery: `WHERE (JSONExtract(JSON_VALUE(body, '$."status"'), 'Float64') > ? AND JSON_EXISTS(body, '$."status"'))`,
expectedArgs: []any{float64(200)},
expectedQuery: `WHERE (JSONExtract(JSON_VALUE(body, '$."status"'), 'Float64') > ? AND (JSON_EXISTS(body, '$."status"') AND LOWER(body) LIKE LOWER(?)))`,
expectedArgs: []any{float64(200), "%\"status\"%"},
expectedErrorContains: "",
},
{
category: "json",
query: "body.message REGEXP 'a*'",
shouldPass: true,
expectedQuery: `WHERE (match(JSON_VALUE(body, '$."message"'), ?) AND JSON_EXISTS(body, '$."message"'))`,
expectedArgs: []any{"a*"},
expectedQuery: `WHERE (match(JSON_VALUE(body, '$."message"'), ?) AND (JSON_EXISTS(body, '$."message"') AND LOWER(body) LIKE LOWER(?)))`,
expectedArgs: []any{"a*", "%\"message\"%"},
expectedErrorContains: "",
},
{
category: "json",
query: `body.message CONTAINS "hello 'world'"`,
shouldPass: true,
expectedQuery: `WHERE (LOWER(JSON_VALUE(body, '$."message"')) LIKE LOWER(?) AND JSON_EXISTS(body, '$."message"'))`,
expectedArgs: []any{"%hello 'world'%"},
expectedQuery: `WHERE ((LOWER(JSON_VALUE(body, '$."message"')) LIKE LOWER(?) AND LOWER(body) LIKE LOWER(?)) AND (JSON_EXISTS(body, '$."message"') AND LOWER(body) LIKE LOWER(?)))`,
expectedArgs: []any{"%hello 'world'%", "%hello 'world'%", "%\"message\"%"},
expectedErrorContains: "",
},
{
category: "json",
query: `body.message EXISTS`,
shouldPass: true,
expectedQuery: `WHERE JSON_EXISTS(body, '$."message"')`,
expectedQuery: `WHERE (JSON_EXISTS(body, '$."message"') AND LOWER(body) LIKE LOWER(?))`,
expectedArgs: []any{"%\"message\"%"},
expectedErrorContains: "",
},
{
category: "json",
query: `body.name IN ('hello', 'world')`,
shouldPass: true,
expectedQuery: `WHERE ((JSON_VALUE(body, '$."name"') = ? OR JSON_VALUE(body, '$."name"') = ?) AND JSON_EXISTS(body, '$."name"'))`,
expectedArgs: []any{"hello", "world"},
expectedQuery: `WHERE (((JSON_VALUE(body, '$."name"') = ? AND LOWER(body) LIKE LOWER(?)) OR (JSON_VALUE(body, '$."name"') = ? AND LOWER(body) LIKE LOWER(?))) AND (JSON_EXISTS(body, '$."name"') AND LOWER(body) LIKE LOWER(?)))`,
expectedArgs: []any{"hello", "%hello%", "world", "%world%", "%\"name\"%"},
expectedErrorContains: "",
},
{
category: "json",
query: `body.value IN (200, 300)`,
shouldPass: true,
expectedQuery: `WHERE ((JSONExtract(JSON_VALUE(body, '$."value"'), 'Float64') = ? OR JSONExtract(JSON_VALUE(body, '$."value"'), 'Float64') = ?) AND JSON_EXISTS(body, '$."value"'))`,
expectedArgs: []any{float64(200), float64(300)},
expectedQuery: `WHERE ((JSONExtract(JSON_VALUE(body, '$."value"'), 'Float64') = ? OR JSONExtract(JSON_VALUE(body, '$."value"'), 'Float64') = ?) AND (JSON_EXISTS(body, '$."value"') AND LOWER(body) LIKE LOWER(?)))`,
expectedArgs: []any{float64(200), float64(300), "%\"value\"%"},
expectedErrorContains: "",
},
{
category: "json",
query: "body.key-with-hyphen = true",
shouldPass: true,
expectedQuery: `WHERE (JSONExtract(JSON_VALUE(body, '$."key-with-hyphen"'), 'Bool') = ? AND JSON_EXISTS(body, '$."key-with-hyphen"'))`,
expectedArgs: []any{true},
expectedQuery: `WHERE (JSONExtract(JSON_VALUE(body, '$."key-with-hyphen"'), 'Bool') = ? AND (JSON_EXISTS(body, '$."key-with-hyphen"') AND LOWER(body) LIKE LOWER(?)))`,
expectedArgs: []any{true, "%\"key-with-hyphen\"%"},
expectedErrorContains: "",
},
}

View File

@@ -495,32 +495,32 @@ func TestFilterExprLogs(t *testing.T) {
category: "FREETEXT with parentheses",
query: "error (status.code=500 OR status.code=503)",
shouldPass: true,
expectedQuery: "WHERE (match(LOWER(body), LOWER(?)) AND (((toFloat64(attributes_number['status.code']) = ? AND mapContains(attributes_number, 'status.code')) OR (toFloat64(attributes_number['status.code']) = ? AND mapContains(attributes_number, 'status.code')))))",
expectedArgs: []any{"error", float64(500), float64(503)},
expectedQuery: "WHERE (match(LOWER(body), LOWER(?)) AND ((((toFloat64(attributes_number['status.code']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status.code')) OR ((toFloat64(attributes_number['status.code']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status.code')))))",
expectedArgs: []any{"error", float64(500), float64(500), float64(503), float64(503)},
expectedErrorContains: "",
},
{
category: "FREETEXT with parentheses",
query: "(status.code=500 OR status.code=503) error",
shouldPass: true,
expectedQuery: "WHERE ((((toFloat64(attributes_number['status.code']) = ? AND mapContains(attributes_number, 'status.code')) OR (toFloat64(attributes_number['status.code']) = ? AND mapContains(attributes_number, 'status.code')))) AND match(LOWER(body), LOWER(?)))",
expectedArgs: []any{float64(500), float64(503), "error"},
expectedQuery: "WHERE (((((toFloat64(attributes_number['status.code']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status.code')) OR ((toFloat64(attributes_number['status.code']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status.code')))) AND match(LOWER(body), LOWER(?)))",
expectedArgs: []any{float64(500), float64(500), float64(503), float64(503), "error"},
expectedErrorContains: "",
},
{
category: "FREETEXT with parentheses",
query: "error AND (status.code=500 OR status.code=503)",
shouldPass: true,
expectedQuery: "WHERE (match(LOWER(body), LOWER(?)) AND (((toFloat64(attributes_number['status.code']) = ? AND mapContains(attributes_number, 'status.code')) OR (toFloat64(attributes_number['status.code']) = ? AND mapContains(attributes_number, 'status.code')))))",
expectedArgs: []any{"error", float64(500), float64(503)},
expectedQuery: "WHERE (match(LOWER(body), LOWER(?)) AND ((((toFloat64(attributes_number['status.code']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status.code')) OR ((toFloat64(attributes_number['status.code']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status.code')))))",
expectedArgs: []any{"error", float64(500), float64(500), float64(503), float64(503)},
expectedErrorContains: "",
},
{
category: "FREETEXT with parentheses",
query: "(status.code=500 OR status.code=503) AND error",
shouldPass: true,
expectedQuery: "WHERE ((((toFloat64(attributes_number['status.code']) = ? AND mapContains(attributes_number, 'status.code')) OR (toFloat64(attributes_number['status.code']) = ? AND mapContains(attributes_number, 'status.code')))) AND match(LOWER(body), LOWER(?)))",
expectedArgs: []any{float64(500), float64(503), "error"},
expectedQuery: "WHERE (((((toFloat64(attributes_number['status.code']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status.code')) OR ((toFloat64(attributes_number['status.code']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status.code')))) AND match(LOWER(body), LOWER(?)))",
expectedArgs: []any{float64(500), float64(500), float64(503), float64(503), "error"},
expectedErrorContains: "",
},
@@ -737,8 +737,8 @@ func TestFilterExprLogs(t *testing.T) {
category: "Key-operator-value boundary",
query: "greater>than",
shouldPass: true,
expectedQuery: `WHERE ((attributes_string['greater'] > ? AND mapContains(attributes_string, 'greater')) OR (JSON_VALUE(body, '$."greater"') > ? AND JSON_EXISTS(body, '$."greater"')))`,
expectedArgs: []any{"than", "than"},
expectedQuery: `WHERE ((attributes_string['greater'] > ? AND mapContains(attributes_string, 'greater')) OR (JSON_VALUE(body, '$."greater"') > ? AND (JSON_EXISTS(body, '$."greater"') AND LOWER(body) LIKE LOWER(?))))`,
expectedArgs: []any{"than", "than", "%\"greater\"%"},
expectedErrorContains: "",
},
{
@@ -753,8 +753,8 @@ func TestFilterExprLogs(t *testing.T) {
category: "Key-operator-value boundary",
query: "less<than",
shouldPass: true,
expectedQuery: `WHERE ((attributes_string['less'] < ? AND mapContains(attributes_string, 'less')) OR (JSON_VALUE(body, '$."less"') < ? AND JSON_EXISTS(body, '$."less"')))`,
expectedArgs: []any{"than", "than"},
expectedQuery: `WHERE ((attributes_string['less'] < ? AND mapContains(attributes_string, 'less')) OR (JSON_VALUE(body, '$."less"') < ? AND (JSON_EXISTS(body, '$."less"') AND LOWER(body) LIKE LOWER(?))))`,
expectedArgs: []any{"than", "than", "%\"less\"%"},
expectedErrorContains: "",
},
{
@@ -809,8 +809,8 @@ func TestFilterExprLogs(t *testing.T) {
category: "Key-operator-value boundary",
query: "user=admin",
shouldPass: true,
expectedQuery: `WHERE ((attributes_string['user'] = ? AND mapContains(attributes_string, 'user')) OR (JSON_VALUE(body, '$."user"') = ? AND JSON_EXISTS(body, '$."user"')))`,
expectedArgs: []any{"admin", "admin"},
expectedQuery: `WHERE ((attributes_string['user'] = ? AND mapContains(attributes_string, 'user')) OR ((JSON_VALUE(body, '$."user"') = ? AND LOWER(body) LIKE LOWER(?)) AND (JSON_EXISTS(body, '$."user"') AND LOWER(body) LIKE LOWER(?))))`,
expectedArgs: []any{"admin", "admin", "%admin%", "%\"user\"%"},
expectedErrorContains: "",
},
{
@@ -827,16 +827,16 @@ func TestFilterExprLogs(t *testing.T) {
category: "Basic equality",
query: "status=200",
shouldPass: true,
expectedQuery: "WHERE (toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status'))",
expectedArgs: []any{float64(200)},
expectedQuery: "WHERE ((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status'))",
expectedArgs: []any{float64(200), float64(200)},
expectedErrorContains: "",
},
{
category: "Basic equality",
query: "code=400",
shouldPass: true,
expectedQuery: "WHERE (toFloat64(attributes_number['code']) = ? AND mapContains(attributes_number, 'code'))",
expectedArgs: []any{float64(400)},
expectedQuery: "WHERE ((toFloat64(attributes_number['code']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'code'))",
expectedArgs: []any{float64(400), float64(400)},
expectedErrorContains: "",
},
{
@@ -867,8 +867,8 @@ func TestFilterExprLogs(t *testing.T) {
category: "Basic equality",
query: "count=0",
shouldPass: true,
expectedQuery: "WHERE (toFloat64(attributes_number['count']) = ? AND mapContains(attributes_number, 'count'))",
expectedArgs: []any{float64(0)},
expectedQuery: "WHERE ((toFloat64(attributes_number['count']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'count'))",
expectedArgs: []any{float64(0), float64(0)},
expectedErrorContains: "",
},
{
@@ -1187,16 +1187,16 @@ func TestFilterExprLogs(t *testing.T) {
category: "IN operator (parentheses)",
query: "status IN (200, 201, 202)",
shouldPass: true,
expectedQuery: "WHERE ((toFloat64(attributes_number['status']) = ? OR toFloat64(attributes_number['status']) = ? OR toFloat64(attributes_number['status']) = ?) AND mapContains(attributes_number, 'status'))",
expectedArgs: []any{float64(200), float64(201), float64(202)},
expectedQuery: "WHERE (((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) OR (toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) OR (toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?))) AND mapContains(attributes_number, 'status'))",
expectedArgs: []any{float64(200), float64(200), float64(201), float64(201), float64(202), float64(202)},
expectedErrorContains: "",
},
{
category: "IN operator (parentheses)",
query: "error.code IN (404, 500, 503)",
shouldPass: true,
expectedQuery: "WHERE ((toFloat64(attributes_number['error.code']) = ? OR toFloat64(attributes_number['error.code']) = ? OR toFloat64(attributes_number['error.code']) = ?) AND mapContains(attributes_number, 'error.code'))",
expectedArgs: []any{float64(404), float64(500), float64(503)},
expectedQuery: "WHERE (((toFloat64(attributes_number['error.code']) = ? AND has(mapValues(attributes_number), ?)) OR (toFloat64(attributes_number['error.code']) = ? AND has(mapValues(attributes_number), ?)) OR (toFloat64(attributes_number['error.code']) = ? AND has(mapValues(attributes_number), ?))) AND mapContains(attributes_number, 'error.code'))",
expectedArgs: []any{float64(404), float64(404), float64(500), float64(500), float64(503), float64(503)},
expectedErrorContains: "",
},
{
@@ -1221,16 +1221,16 @@ func TestFilterExprLogs(t *testing.T) {
category: "IN operator (brackets)",
query: "status IN [200, 201, 202]",
shouldPass: true,
expectedQuery: "WHERE ((toFloat64(attributes_number['status']) = ? OR toFloat64(attributes_number['status']) = ? OR toFloat64(attributes_number['status']) = ?) AND mapContains(attributes_number, 'status'))",
expectedArgs: []any{float64(200), float64(201), float64(202)},
expectedQuery: "WHERE (((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) OR (toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) OR (toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?))) AND mapContains(attributes_number, 'status'))",
expectedArgs: []any{float64(200), float64(200), float64(201), float64(201), float64(202), float64(202)},
expectedErrorContains: "",
},
{
category: "IN operator (brackets)",
query: "error.code IN [404, 500, 503]",
shouldPass: true,
expectedQuery: "WHERE ((toFloat64(attributes_number['error.code']) = ? OR toFloat64(attributes_number['error.code']) = ? OR toFloat64(attributes_number['error.code']) = ?) AND mapContains(attributes_number, 'error.code'))",
expectedArgs: []any{float64(404), float64(500), float64(503)},
expectedQuery: "WHERE (((toFloat64(attributes_number['error.code']) = ? AND has(mapValues(attributes_number), ?)) OR (toFloat64(attributes_number['error.code']) = ? AND has(mapValues(attributes_number), ?)) OR (toFloat64(attributes_number['error.code']) = ? AND has(mapValues(attributes_number), ?))) AND mapContains(attributes_number, 'error.code'))",
expectedArgs: []any{float64(404), float64(404), float64(500), float64(500), float64(503), float64(503)},
expectedErrorContains: "",
},
{
@@ -1561,15 +1561,15 @@ func TestFilterExprLogs(t *testing.T) {
expectedArgs: []any{"download"},
expectedErrorContains: "function `hasToken` expects value parameter to be a string",
},
// A multi-token needle (separator/whitespace) is a clean 400, not a CH execution error.
// A multi-token value (separator/whitespace) is a clean 400, not a CH execution error.
{
category: "hasTokenUnderscoreNeedle",
category: "hasTokenUnderscoreSeparator",
query: "hasToken(body, \"user_id\")",
shouldPass: false,
expectedErrorContains: "function `hasToken` matches a single whole token",
},
{
category: "hasTokenWhitespaceNeedle",
category: "hasTokenWhitespaceSeparator",
query: "hasToken(body, \"production node\")",
shouldPass: false,
expectedErrorContains: "function `hasToken` matches a single whole token",
@@ -1609,8 +1609,8 @@ func TestFilterExprLogs(t *testing.T) {
category: "Explicit AND",
query: "status=200 AND service.name=\"api\"",
shouldPass: true,
expectedQuery: "WHERE ((toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status')) AND (multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL))",
expectedArgs: []any{float64(200), "api"},
expectedQuery: "WHERE (((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status')) AND (multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL))",
expectedArgs: []any{float64(200), float64(200), "api"},
expectedErrorContains: "",
},
{
@@ -1635,8 +1635,8 @@ func TestFilterExprLogs(t *testing.T) {
category: "Explicit OR",
query: "status=200 OR status=201",
shouldPass: true,
expectedQuery: "WHERE ((toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status')) OR (toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status')))",
expectedArgs: []any{float64(200), float64(201)},
expectedQuery: "WHERE (((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status')) OR ((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status')))",
expectedArgs: []any{float64(200), float64(200), float64(201), float64(201)},
expectedErrorContains: "",
},
{
@@ -1661,8 +1661,8 @@ func TestFilterExprLogs(t *testing.T) {
category: "NOT with expressions",
query: "NOT status=200",
shouldPass: true,
expectedQuery: "WHERE NOT ((toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status')))",
expectedArgs: []any{float64(200)},
expectedQuery: "WHERE NOT (((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status')))",
expectedArgs: []any{float64(200), float64(200)},
expectedErrorContains: "",
},
{
@@ -1687,8 +1687,8 @@ func TestFilterExprLogs(t *testing.T) {
category: "AND + OR combinations",
query: "status=200 AND (service.name=\"api\" OR service.name=\"web\")",
shouldPass: true,
expectedQuery: "WHERE ((toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status')) AND (((multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL) OR (multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL))))",
expectedArgs: []any{float64(200), "api", "web"},
expectedQuery: "WHERE (((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status')) AND (((multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL) OR (multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL))))",
expectedArgs: []any{float64(200), float64(200), "api", "web"},
expectedErrorContains: "",
},
{
@@ -1713,8 +1713,8 @@ func TestFilterExprLogs(t *testing.T) {
category: "AND + NOT combinations",
query: "status=200 AND NOT service.name=\"api\"",
shouldPass: true,
expectedQuery: "WHERE ((toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status')) AND NOT ((multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL)))",
expectedArgs: []any{float64(200), "api"},
expectedQuery: "WHERE (((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status')) AND NOT ((multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL)))",
expectedArgs: []any{float64(200), float64(200), "api"},
expectedErrorContains: "",
},
{
@@ -1731,8 +1731,8 @@ func TestFilterExprLogs(t *testing.T) {
category: "OR + NOT combinations",
query: "NOT status=200 OR NOT service.name=\"api\"",
shouldPass: true,
expectedQuery: "WHERE (NOT ((toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status'))) OR NOT ((multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL)))",
expectedArgs: []any{float64(200), "api"},
expectedQuery: "WHERE (NOT (((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status'))) OR NOT ((multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL)))",
expectedArgs: []any{float64(200), float64(200), "api"},
expectedErrorContains: "",
},
{
@@ -1749,8 +1749,8 @@ func TestFilterExprLogs(t *testing.T) {
category: "AND + OR + NOT combinations",
query: "status=200 AND (service.name=\"api\" OR NOT duration>1000)",
shouldPass: true,
expectedQuery: "WHERE ((toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status')) AND (((multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL) OR NOT ((toFloat64(attributes_number['duration']) > ? AND mapContains(attributes_number, 'duration'))))))",
expectedArgs: []any{float64(200), "api", float64(1000)},
expectedQuery: "WHERE (((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status')) AND (((multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL) OR NOT ((toFloat64(attributes_number['duration']) > ? AND mapContains(attributes_number, 'duration'))))))",
expectedArgs: []any{float64(200), float64(200), "api", float64(1000)},
expectedErrorContains: "",
},
{
@@ -1765,8 +1765,8 @@ func TestFilterExprLogs(t *testing.T) {
category: "AND + OR + NOT combinations",
query: "NOT (status=200 AND service.name=\"api\") OR count>0",
shouldPass: true,
expectedQuery: "WHERE (NOT ((((toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status')) AND (multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL)))) OR (toFloat64(attributes_number['count']) > ? AND mapContains(attributes_number, 'count')))",
expectedArgs: []any{float64(200), "api", float64(0)},
expectedQuery: "WHERE (NOT (((((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status')) AND (multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL)))) OR (toFloat64(attributes_number['count']) > ? AND mapContains(attributes_number, 'count')))",
expectedArgs: []any{float64(200), float64(200), "api", float64(0)},
expectedErrorContains: "",
},
@@ -1775,8 +1775,8 @@ func TestFilterExprLogs(t *testing.T) {
category: "Implicit AND",
query: "status=200 service.name=\"api\"",
shouldPass: true,
expectedQuery: "WHERE ((toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status')) AND (multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL))",
expectedArgs: []any{float64(200), "api"},
expectedQuery: "WHERE (((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status')) AND (multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL))",
expectedArgs: []any{float64(200), float64(200), "api"},
expectedErrorContains: "",
},
{
@@ -1801,8 +1801,8 @@ func TestFilterExprLogs(t *testing.T) {
category: "Mixed implicit/explicit AND",
query: "status=200 AND service.name=\"api\" duration<1000",
shouldPass: true,
expectedQuery: "WHERE ((toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status')) AND (multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL) AND (toFloat64(attributes_number['duration']) < ? AND mapContains(attributes_number, 'duration')))",
expectedArgs: []any{float64(200), "api", float64(1000)},
expectedQuery: "WHERE (((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status')) AND (multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL) AND (toFloat64(attributes_number['duration']) < ? AND mapContains(attributes_number, 'duration')))",
expectedArgs: []any{float64(200), float64(200), "api", float64(1000)},
expectedErrorContains: "",
},
{
@@ -1819,8 +1819,8 @@ func TestFilterExprLogs(t *testing.T) {
category: "Simple grouping",
query: "(status=200)",
shouldPass: true,
expectedQuery: "WHERE ((toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status')))",
expectedArgs: []any{float64(200)},
expectedQuery: "WHERE (((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status')))",
expectedArgs: []any{float64(200), float64(200)},
expectedErrorContains: "",
},
{
@@ -1845,8 +1845,8 @@ func TestFilterExprLogs(t *testing.T) {
category: "Nested grouping",
query: "((status=200))",
shouldPass: true,
expectedQuery: "WHERE (((toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status'))))",
expectedArgs: []any{float64(200)},
expectedQuery: "WHERE ((((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status'))))",
expectedArgs: []any{float64(200), float64(200)},
expectedErrorContains: "",
},
{
@@ -1871,8 +1871,8 @@ func TestFilterExprLogs(t *testing.T) {
category: "Complex nested grouping",
query: "(status=200 AND (service.name=\"api\" OR service.name=\"web\"))",
shouldPass: true,
expectedQuery: "WHERE (((toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status')) AND (((multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL) OR (multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL)))))",
expectedArgs: []any{float64(200), "api", "web"},
expectedQuery: "WHERE ((((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status')) AND (((multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL) OR (multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL)))))",
expectedArgs: []any{float64(200), float64(200), "api", "web"},
expectedErrorContains: "",
},
{
@@ -1897,8 +1897,8 @@ func TestFilterExprLogs(t *testing.T) {
category: "Deep nesting",
query: "(((status=200 OR status=201) AND service.name=\"api\") OR ((status=202 OR status=203) AND service.name=\"web\"))",
shouldPass: true,
expectedQuery: "WHERE (((((((toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status')) OR (toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status')))) AND (multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL))) OR (((((toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status')) OR (toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status')))) AND (multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL)))))",
expectedArgs: []any{float64(200), float64(201), "api", float64(202), float64(203), "web"},
expectedQuery: "WHERE ((((((((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status')) OR ((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status')))) AND (multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL))) OR ((((((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status')) OR ((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status')))) AND (multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL)))))",
expectedArgs: []any{float64(200), float64(200), float64(201), float64(201), "api", float64(202), float64(202), float64(203), float64(203), "web"},
expectedErrorContains: "",
},
{
@@ -1949,32 +1949,32 @@ func TestFilterExprLogs(t *testing.T) {
category: "Numeric values",
query: "status=200",
shouldPass: true,
expectedQuery: "WHERE (toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status'))",
expectedArgs: []any{float64(200)},
expectedQuery: "WHERE ((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status'))",
expectedArgs: []any{float64(200), float64(200)},
expectedErrorContains: "",
},
{
category: "Numeric values",
query: "count=0",
shouldPass: true,
expectedQuery: "WHERE (toFloat64(attributes_number['count']) = ? AND mapContains(attributes_number, 'count'))",
expectedArgs: []any{float64(0)},
expectedQuery: "WHERE ((toFloat64(attributes_number['count']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'count'))",
expectedArgs: []any{float64(0), float64(0)},
expectedErrorContains: "",
},
{
category: "Numeric values",
query: "duration=1000.5",
shouldPass: true,
expectedQuery: "WHERE (toFloat64(attributes_number['duration']) = ? AND mapContains(attributes_number, 'duration'))",
expectedArgs: []any{float64(1000.5)},
expectedQuery: "WHERE ((toFloat64(attributes_number['duration']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'duration'))",
expectedArgs: []any{float64(1000.5), float64(1000.5)},
expectedErrorContains: "",
},
{
category: "Numeric values",
query: "amount=-10.25",
shouldPass: true,
expectedQuery: "WHERE (toFloat64(attributes_number['amount']) = ? AND mapContains(attributes_number, 'amount'))",
expectedArgs: []any{float64(-10.25)},
expectedQuery: "WHERE ((toFloat64(attributes_number['amount']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'amount'))",
expectedArgs: []any{float64(-10.25), float64(-10.25)},
expectedErrorContains: "",
},
@@ -2052,8 +2052,8 @@ func TestFilterExprLogs(t *testing.T) {
category: "Nested object paths",
query: "response.body.data.items[].id=123",
shouldPass: true,
expectedQuery: `WHERE ((toFloat64(attributes_number['response.body.data.items[].id']) = ? AND mapContains(attributes_number, 'response.body.data.items[].id')) OR (JSONExtract(JSON_VALUE(body, '$."response"."body"."data"."items"[*]."id"'), 'Float64') = ? AND JSON_EXISTS(body, '$."response"."body"."data"."items"[*]."id"')))`,
expectedArgs: []any{float64(123), float64(123)},
expectedQuery: `WHERE (((toFloat64(attributes_number['response.body.data.items[].id']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'response.body.data.items[].id')) OR (JSONExtract(JSON_VALUE(body, '$."response"."body"."data"."items"[*]."id"'), 'Float64') = ? AND (JSON_EXISTS(body, '$."response"."body"."data"."items"[*]."id"') AND LOWER(body) LIKE LOWER(?))))`,
expectedArgs: []any{float64(123), float64(123), float64(123), "%\"response\"%\"body\"%\"data\"%\"items\"%\"id\"%"},
expectedErrorContains: "",
},
{
@@ -2083,29 +2083,29 @@ func TestFilterExprLogs(t *testing.T) {
category: "Operator precedence",
query: "NOT status=200 AND service.name=\"api\"",
shouldPass: true,
expectedQuery: "WHERE (NOT ((toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status'))) AND (multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL))",
expectedArgs: []any{float64(200), "api"}, // Should be (NOT status=200) AND service.name="api"
expectedQuery: "WHERE (NOT (((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status'))) AND (multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL))",
expectedArgs: []any{float64(200), float64(200), "api"}, // Should be (NOT status=200) AND service.name="api"
},
{
category: "Operator precedence",
query: "status=200 AND service.name=\"api\" OR service.name=\"web\"",
shouldPass: true,
expectedQuery: "WHERE (((toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status')) AND (multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL)) OR (multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL))",
expectedArgs: []any{float64(200), "api", "web"}, // Should be (status=200 AND service.name="api") OR service.name="web"
expectedQuery: "WHERE ((((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status')) AND (multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL)) OR (multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL))",
expectedArgs: []any{float64(200), float64(200), "api", "web"}, // Should be (status=200 AND service.name="api") OR service.name="web"
},
{
category: "Operator precedence",
query: "NOT status=200 OR NOT service.name=\"api\"",
shouldPass: true,
expectedQuery: "WHERE (NOT ((toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status'))) OR NOT ((multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL)))",
expectedArgs: []any{float64(200), "api"}, // Should be (NOT status=200) OR (NOT service.name="api")
expectedQuery: "WHERE (NOT (((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status'))) OR NOT ((multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL)))",
expectedArgs: []any{float64(200), float64(200), "api"}, // Should be (NOT status=200) OR (NOT service.name="api")
},
{
category: "Operator precedence",
query: "status=200 OR service.name=\"api\" AND level=\"ERROR\"",
shouldPass: true,
expectedQuery: "WHERE ((toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status')) OR ((multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL) AND (attributes_string['level'] = ? AND mapContains(attributes_string, 'level'))))",
expectedArgs: []any{float64(200), "api", "ERROR"}, // Should be status=200 OR (service.name="api" AND level="ERROR")
expectedQuery: "WHERE (((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status')) OR ((multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL) AND (attributes_string['level'] = ? AND mapContains(attributes_string, 'level'))))",
expectedArgs: []any{float64(200), float64(200), "api", "ERROR"}, // Should be status=200 OR (service.name="api" AND level="ERROR")
},
// Different whitespace patterns
@@ -2129,8 +2129,8 @@ func TestFilterExprLogs(t *testing.T) {
category: "Whitespace patterns",
query: "status=200 AND service.name=\"api\"",
shouldPass: true,
expectedQuery: "WHERE ((toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status')) AND (multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL))",
expectedArgs: []any{float64(200), "api"}, // Multiple spaces
expectedQuery: "WHERE (((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status')) AND (multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL))",
expectedArgs: []any{float64(200), float64(200), "api"}, // Multiple spaces
},
// More Unicode characters
@@ -2365,8 +2365,8 @@ func TestFilterExprLogs(t *testing.T) {
category: "Unusual whitespace",
query: "status = 200 AND service.name = \"api\"",
shouldPass: true,
expectedQuery: "WHERE ((toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status')) AND (multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL))",
expectedArgs: []any{float64(200), "api"},
expectedQuery: "WHERE (((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status')) AND (multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL))",
expectedArgs: []any{float64(200), float64(200), "api"},
},
{
category: "Unusual whitespace",
@@ -2426,9 +2426,9 @@ func TestFilterExprLogs(t *testing.T) {
)
`,
shouldPass: true,
expectedQuery: "WHERE ((((((((toFloat64(attributes_number['status']) >= ? AND mapContains(attributes_number, 'status')) AND (toFloat64(attributes_number['status']) < ? AND mapContains(attributes_number, 'status')))) OR (((toFloat64(attributes_number['status']) >= ? AND mapContains(attributes_number, 'status')) AND (toFloat64(attributes_number['status']) < ? AND mapContains(attributes_number, 'status')) AND NOT ((toFloat64(attributes_number['status']) = ? AND mapContains(attributes_number, 'status'))))))) AND ((((multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? OR multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? OR multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ?) AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL) OR (((multiIf(resource.`service.type` IS NOT NULL, resource.`service.type`::String, mapContains(resources_string, 'service.type'), resources_string['service.type'], NULL) = ? AND multiIf(resource.`service.type` IS NOT NULL, resource.`service.type`::String, mapContains(resources_string, 'service.type'), resources_string['service.type'], NULL) IS NOT NULL) AND NOT ((multiIf(resource.`service.deprecated` IS NOT NULL, resource.`service.deprecated`::String, mapContains(resources_string, 'service.deprecated'), resources_string['service.deprecated'], NULL) = ? AND multiIf(resource.`service.deprecated` IS NOT NULL, resource.`service.deprecated`::String, mapContains(resources_string, 'service.deprecated'), resources_string['service.deprecated'], NULL) IS NOT NULL)))))))) AND (((((toFloat64(attributes_number['duration']) < ? AND mapContains(attributes_number, 'duration')) OR ((toFloat64(attributes_number['duration']) BETWEEN ? AND ? AND mapContains(attributes_number, 'duration'))))) AND ((multiIf(resource.`environment` IS NOT NULL, resource.`environment`::String, mapContains(resources_string, 'environment'), resources_string['environment'], NULL) <> ? OR (((multiIf(resource.`environment` IS NOT NULL, resource.`environment`::String, mapContains(resources_string, 'environment'), resources_string['environment'], NULL) = ? AND multiIf(resource.`environment` IS NOT NULL, resource.`environment`::String, mapContains(resources_string, 'environment'), resources_string['environment'], NULL) IS NOT NULL) AND (attributes_bool['is_automated_test'] = ? AND mapContains(attributes_bool, 'is_automated_test')))))))) AND NOT ((((((LOWER(attributes_string['message']) LIKE LOWER(?) AND mapContains(attributes_string, 'message')) OR (LOWER(attributes_string['message']) LIKE LOWER(?) AND mapContains(attributes_string, 'message')))) AND (attributes_string['severity'] = ? AND mapContains(attributes_string, 'severity'))))))",
expectedQuery: "WHERE ((((((((toFloat64(attributes_number['status']) >= ? AND mapContains(attributes_number, 'status')) AND (toFloat64(attributes_number['status']) < ? AND mapContains(attributes_number, 'status')))) OR (((toFloat64(attributes_number['status']) >= ? AND mapContains(attributes_number, 'status')) AND (toFloat64(attributes_number['status']) < ? AND mapContains(attributes_number, 'status')) AND NOT (((toFloat64(attributes_number['status']) = ? AND has(mapValues(attributes_number), ?)) AND mapContains(attributes_number, 'status'))))))) AND ((((multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? OR multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ? OR multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) = ?) AND multiIf(resource.`service.name` IS NOT NULL, resource.`service.name`::String, mapContains(resources_string, 'service.name'), resources_string['service.name'], NULL) IS NOT NULL) OR (((multiIf(resource.`service.type` IS NOT NULL, resource.`service.type`::String, mapContains(resources_string, 'service.type'), resources_string['service.type'], NULL) = ? AND multiIf(resource.`service.type` IS NOT NULL, resource.`service.type`::String, mapContains(resources_string, 'service.type'), resources_string['service.type'], NULL) IS NOT NULL) AND NOT ((multiIf(resource.`service.deprecated` IS NOT NULL, resource.`service.deprecated`::String, mapContains(resources_string, 'service.deprecated'), resources_string['service.deprecated'], NULL) = ? AND multiIf(resource.`service.deprecated` IS NOT NULL, resource.`service.deprecated`::String, mapContains(resources_string, 'service.deprecated'), resources_string['service.deprecated'], NULL) IS NOT NULL)))))))) AND (((((toFloat64(attributes_number['duration']) < ? AND mapContains(attributes_number, 'duration')) OR ((toFloat64(attributes_number['duration']) BETWEEN ? AND ? AND mapContains(attributes_number, 'duration'))))) AND ((multiIf(resource.`environment` IS NOT NULL, resource.`environment`::String, mapContains(resources_string, 'environment'), resources_string['environment'], NULL) <> ? OR (((multiIf(resource.`environment` IS NOT NULL, resource.`environment`::String, mapContains(resources_string, 'environment'), resources_string['environment'], NULL) = ? AND multiIf(resource.`environment` IS NOT NULL, resource.`environment`::String, mapContains(resources_string, 'environment'), resources_string['environment'], NULL) IS NOT NULL) AND (attributes_bool['is_automated_test'] = ? AND mapContains(attributes_bool, 'is_automated_test')))))))) AND NOT ((((((LOWER(attributes_string['message']) LIKE LOWER(?) AND mapContains(attributes_string, 'message')) OR (LOWER(attributes_string['message']) LIKE LOWER(?) AND mapContains(attributes_string, 'message')))) AND (attributes_string['severity'] = ? AND mapContains(attributes_string, 'severity'))))))",
expectedArgs: []any{
float64(200), float64(300), float64(400), float64(500), float64(404),
float64(200), float64(300), float64(400), float64(500), float64(404), float64(404),
"api", "web", "auth",
"internal", true,
float64(1000), float64(1000), float64(5000),
@@ -2521,7 +2521,7 @@ func TestFilterExprLogsConflictNegation(t *testing.T) {
query: "body NOT LIKE 'done'",
shouldPass: true,
// lower index search on body even for LIKE
expectedQuery: "WHERE (LOWER(body) NOT LIKE LOWER(?) AND attributes_string['body'] NOT LIKE ?)",
expectedQuery: "WHERE (body NOT LIKE ? AND attributes_string['body'] NOT LIKE ?)",
expectedArgs: []any{"done", "done"},
expectedErrorContains: "",
},

View File

@@ -439,27 +439,27 @@ func (c *jsonConditionBuilder) arrayFuncScalarLeaf(node *telemetrytypes.JSONAcce
// buildTokenFunctionCondition builds a hasToken search over a body JSON string field:
// hasToken(LOWER(<elem>), LOWER(?)) wrapped in arrayExists over any array hops between the
// root and the terminal. The field must resolve to a String leaf or a String array.
func (c *jsonConditionBuilder) buildTokenFunctionCondition(needle any, sb *sqlbuilder.SelectBuilder) (string, error) {
func (c *jsonConditionBuilder) buildTokenFunctionCondition(token any, sb *sqlbuilder.SelectBuilder) (string, error) {
if len(c.key.JSONPlan) == 0 {
return "", errors.NewInvalidInputf(errors.CodeInvalidInput, "function `hasToken` could not resolve a JSON access plan for field `%s`", c.key.Name)
}
return c.buildOredRootChains(func(node *telemetrytypes.JSONAccessNode) (string, error) {
return c.tokenLeaf(node, needle, sb)
return c.tokenLeaf(node, token, sb)
}, sb)
}
// tokenLeaf builds the hasToken match at a terminal node: a direct match for a String leaf
// (coalesced to false, as in arrayFuncScalarLeaf), or an arrayExists over the elements for a
// String array leaf. hasToken is string-only, so any other element type is rejected.
func (c *jsonConditionBuilder) tokenLeaf(node *telemetrytypes.JSONAccessNode, needle any, sb *sqlbuilder.SelectBuilder) (string, error) {
func (c *jsonConditionBuilder) tokenLeaf(node *telemetrytypes.JSONAccessNode, token any, sb *sqlbuilder.SelectBuilder) (string, error) {
switch node.TerminalConfig.ElemType {
case telemetrytypes.String:
fieldExpr := fmt.Sprintf("dynamicElement(%s, 'String')", node.FieldPath())
return fmt.Sprintf("ifNull(hasToken(LOWER(%s), LOWER(%s)), false)", fieldExpr, sb.Var(needle)), nil
return fmt.Sprintf("ifNull(hasToken(LOWER(%s), LOWER(%s)), false)", fieldExpr, sb.Var(token)), nil
case telemetrytypes.ArrayString:
arrayExpr := fmt.Sprintf("dynamicElement(%s, '%s')", node.FieldPath(), node.TerminalConfig.ElemType.StringValue())
return fmt.Sprintf("arrayExists(x -> hasToken(LOWER(x), LOWER(%s)), %s)", sb.Var(needle), arrayExpr), nil
return fmt.Sprintf("arrayExists(x -> hasToken(LOWER(x), LOWER(%s)), %s)", sb.Var(token), arrayExpr), nil
default:
return "", errors.NewInvalidInputf(errors.CodeInvalidInput, "function `hasToken` only supports string fields; field `%s` is `%s`", c.key.Name, node.TerminalConfig.Key.FieldDataType.StringValue())
}

View File

@@ -9,6 +9,8 @@ import (
qbtypes "github.com/SigNoz/signoz/pkg/types/querybuildertypes/querybuildertypesv5"
"github.com/SigNoz/signoz/pkg/types/telemetrytypes"
"github.com/huandu/go-sqlbuilder"
)
func parseStrValue(valueStr string, operator qbtypes.FilterOperator) (telemetrytypes.FieldDataType, any) {
@@ -92,6 +94,114 @@ func InferDataType(value any, operator qbtypes.FilterOperator, key *telemetrytyp
return closure(value, key)
}
// likePatternLiterals returns the runs of pattern between unescaped wildcards, with `\` escapes
// resolved; every value the pattern matches holds each run verbatim. ClickHouse treats `\` as an
// escape only before `%`, `_` and itself, so dropping it elsewhere would yield a run it never requires.
func likePatternLiterals(pattern string) []string {
var (
literals []string
run strings.Builder
)
for i := 0; i < len(pattern); i++ {
switch c := pattern[i]; c {
case '%', '_':
if run.Len() > 0 {
literals = append(literals, run.String())
run.Reset()
}
case '\\':
if i+1 >= len(pattern) {
run.WriteByte('\\')
continue
}
i++
if escaped := pattern[i]; escaped != '%' && escaped != '_' && escaped != '\\' {
run.WriteByte('\\')
}
run.WriteByte(pattern[i])
default:
run.WriteByte(c)
}
}
if run.Len() > 0 {
literals = append(literals, run.String())
}
return literals
}
// jsonEscapable reports whether a JSON encoder is free to rewrite r: `"` and `\` always, `/` by
// PHP, `<` `>` `&` by Go, non-printable ASCII by Python's ensure_ascii. The legacy body holds the
// producer's own text, so a literal spanning one of these may not be there to find.
func jsonEscapable(r rune) bool {
return r < 0x20 || r > 0x7e || strings.ContainsRune(`"\/<>&`, r)
}
// jsonTextRuns splits s at every byte an encoder may rewrite. A body whose JSON holds s contains
// each returned run verbatim, in order.
func jsonTextRuns(s string) []string {
return strings.FieldsFunc(s, jsonEscapable)
}
// bodyPathLiterals returns one literal per component of key's JSON path, quoted the way JSON writes
// an object key. A component holding a byte an encoder may rewrite is dropped; JSON writes a parent
// first, so the order carries.
func bodyPathLiterals(key *telemetrytypes.TelemetryFieldKey) []string {
var literals []string
for _, part := range strings.Split(key.Name, ".") {
if idx := strings.Index(part, "["); idx >= 0 {
part = part[:idx]
}
if literal := `"` + part + `"`; !strings.ContainsFunc(part, jsonEscapable) {
literals = append(literals, literal)
}
}
return literals
}
// bodyValueLiterals returns the literals a comparison implies in the body text. Only string
// comparisons qualify: a number is compared after JSONExtract parses it, which reads 1.23e2
// as 123, so the digits of the filter value need not appear in the body at all.
func bodyValueLiterals(operator qbtypes.FilterOperator, value any) []string {
str, ok := value.(string)
if !ok {
return nil
}
switch operator {
case qbtypes.FilterOperatorEqual, qbtypes.FilterOperatorContains:
return jsonTextRuns(str)
case qbtypes.FilterOperatorLike, qbtypes.FilterOperatorILike:
var literals []string
for _, literal := range likePatternLiterals(str) {
literals = append(literals, jsonTextRuns(literal)...)
}
return literals
}
return nil
}
// escapeLikeLiteral escapes the LIKE metacharacters so s matches as literal text. Backslash
// goes first, being the escape character itself.
func escapeLikeLiteral(s string) string {
s = strings.ReplaceAll(s, `\`, `\\`)
s = strings.ReplaceAll(s, "%", `\%`)
return strings.ReplaceAll(s, "_", `\_`)
}
// bodyIndexPredicate asserts the raw body text holds the literals in order. ILike renders as
// LOWER(body) LIKE LOWER(?) on the ClickHouse flavor — the expression both bloom filters index.
// They are plain text and backslash-free by construction, so only the LIKE wildcards need escaping.
func bodyIndexPredicate(literals []string, sb *sqlbuilder.SelectBuilder) string {
if len(literals) == 0 {
return ""
}
escaped := make([]string, 0, len(literals))
for _, literal := range literals {
escaped = append(escaped, escapeLikeLiteral(literal))
}
pattern := "%" + strings.Join(escaped, "%") + "%"
return sb.ILike(LogsV2BodyColumn, pattern)
}
func getBodyJSONPath(key *telemetrytypes.TelemetryFieldKey) string {
parts := strings.Split(key.Name, ".")
newParts := []string{}
@@ -139,14 +249,14 @@ func GetBodyJSONKeyForExists(_ context.Context, key *telemetrytypes.TelemetryFie
return fmt.Sprintf("JSON_EXISTS(body, '$.%s')", getBodyJSONPath(key))
}
// legacyElemType infers the has-family element type from the needle (legacy has no schema). It
// scans EVERY value so the chosen array type and all coerced needles agree — else ClickHouse
// legacyElemType infers the has-family element type from the arg (legacy has no schema). It
// scans EVERY value so the chosen array type and all coerced args agree — else ClickHouse
// raises "no supertype ... String" (code 386). Int64 stays distinct from Float64 so a quoted
// integer is exact past 2^53 (unquoted literals already arrive as float64, parsed upstream).
func legacyElemType(needle any) telemetrytypes.FieldDataType {
list, ok := needle.([]any)
func legacyElemType(arg any) telemetrytypes.FieldDataType {
list, ok := arg.([]any)
if !ok {
list = []any{needle}
list = []any{arg}
}
if len(list) == 0 {
return telemetrytypes.FieldDataTypeString
@@ -167,7 +277,7 @@ func legacyElemType(needle any) telemetrytypes.FieldDataType {
}
default:
// booleans (and anything else) -> String; a bool renders to 'true'/'false', so a
// bool needle only matches genuine JSON booleans, not truthy numbers/strings.
// bool arg only matches genuine JSON booleans, not truthy numbers/strings.
allInt, allNumeric = false, false
}
}
@@ -181,9 +291,9 @@ func legacyElemType(needle any) telemetrytypes.FieldDataType {
}
}
// legacyCoerceNeedle coerces a needle to elem type dt so its bound-arg type matches the
// legacyCoerceElement coerces an element to elem type dt so its bound-arg type matches the
// extracted column (legacyElemType guarantees it's coercible).
func legacyCoerceNeedle(v any, dt telemetrytypes.FieldDataType) any {
func legacyCoerceElement(v any, dt telemetrytypes.FieldDataType) any {
switch dt {
case telemetrytypes.FieldDataTypeInt64:
if s, ok := v.(string); ok {
@@ -199,7 +309,7 @@ func legacyCoerceNeedle(v any, dt telemetrytypes.FieldDataType) any {
}
return v
default:
return bodyArrayNeedleString(v)
return bodyArrayElementString(v)
}
}
@@ -216,8 +326,8 @@ func getBodyJSONArrayKey(key *telemetrytypes.TelemetryFieldKey, dt telemetrytype
// getBodyJSONScalarKey builds the single-element-set fallback for a scalar body value: the leaf
// extracted as a scalar of type dt, plus a guard restricting it to a genuinely scalar body. The
// guard is required because JSON_VALUE returns '' for an array/object/missing value, which would
// otherwise zero-value match (has(x,0) / has(x,false) / has(x,'') on any array). ok=false when
// guard is required because JSON_VALUE returns for an array/object/missing value, which would
// otherwise zero-value match (has(x,0) / has(x,false) / has(x,) on any array). ok=false when
// the path still traverses an array ([*]/[]).
func getBodyJSONScalarKey(key *telemetrytypes.TelemetryFieldKey, dt telemetrytypes.FieldDataType) (expr string, guard string, ok bool) {
name := strings.TrimSuffix(strings.TrimSuffix(key.Name, "[*]"), "[]")
@@ -242,7 +352,7 @@ func getBodyJSONScalarKey(key *telemetrytypes.TelemetryFieldKey, dt telemetrytyp
return expr, guard, true
}
func bodyArrayNeedleString(v any) string {
func bodyArrayElementString(v any) string {
switch t := v.(type) {
case string:
return t

View File

@@ -391,96 +391,6 @@ func TestConditionForResourceWithEvolution(t *testing.T) {
}
}
// TestConditionForScopeIntrinsicFields covers the scope.name/scope.version intrinsic
// fields against the "scope" JSON column. These are *declared* String paths on that
// column, so a row without a scope reads as ” and never NULL: presence must be an
// empty-string check, since "IS NOT NULL" would hold for every row. That also rules
// out treating them as nested attribute keys under scope.attributes, which are
// undeclared (Dynamic) paths and genuinely NULL when absent.
func TestConditionForScopeIntrinsicFields(t *testing.T) {
ctx := context.Background()
fm := NewFieldMapper(flaggertest.New(t))
conditionBuilder := NewConditionBuilder(fm, flaggertest.New(t))
testCases := []struct {
name string
key telemetrytypes.TelemetryFieldKey
operator qbtypes.FilterOperator
value any
expectedSQL string
}{
{
name: "Equal - scope.name",
key: telemetrytypes.TelemetryFieldKey{
Name: "scope.name",
FieldContext: telemetrytypes.FieldContextScope,
FieldDataType: telemetrytypes.FieldDataTypeString,
},
operator: qbtypes.FilterOperatorEqual,
value: "io.signoz.payment",
expectedSQL: "(scope.name::String = ? AND scope.name::String <> '')",
},
{
name: "Equal - scope.version",
key: telemetrytypes.TelemetryFieldKey{
Name: "scope.version",
FieldContext: telemetrytypes.FieldContextScope,
FieldDataType: telemetrytypes.FieldDataTypeString,
},
operator: qbtypes.FilterOperatorEqual,
value: "2.3.1",
expectedSQL: "(scope.version::String = ? AND scope.version::String <> '')",
},
{
name: "Exists - scope.name",
key: telemetrytypes.TelemetryFieldKey{
Name: "scope.name",
FieldContext: telemetrytypes.FieldContextScope,
FieldDataType: telemetrytypes.FieldDataTypeString,
},
operator: qbtypes.FilterOperatorExists,
value: nil,
expectedSQL: "scope.name::String <> ''",
},
{
name: "NotExists - scope.version",
key: telemetrytypes.TelemetryFieldKey{
Name: "scope.version",
FieldContext: telemetrytypes.FieldContextScope,
FieldDataType: telemetrytypes.FieldDataTypeString,
},
operator: qbtypes.FilterOperatorNotExists,
value: nil,
expectedSQL: "scope.version::String = ''",
},
{
// `scope.attribute.name` (normalized to {attribute.name, scope}) addresses the scope
// attribute named `name` — the declared `scope.name` path is never reached this way.
name: "Equal - scope.attribute.name reaches the named scope attribute",
key: telemetrytypes.TelemetryFieldKey{
Name: "attribute.name",
FieldContext: telemetrytypes.FieldContextScope,
FieldDataType: telemetrytypes.FieldDataTypeString,
},
operator: qbtypes.FilterOperatorEqual,
value: "io.signoz.checkout",
expectedSQL: "(scope.attributes.`name`::String = ? AND scope.attributes.`name` IS NOT NULL)",
},
}
for _, tc := range testCases {
sb := sqlbuilder.NewSelectBuilder()
t.Run(tc.name, func(t *testing.T) {
conds, _, err := conditionBuilder.ConditionFor(ctx, valuer.UUID{}, 0, 0, &tc.key, map[string][]*telemetrytypes.TelemetryFieldKey{tc.key.Name: {&tc.key}}, qbtypes.ConditionBuilderOptions{}, tc.operator, tc.value, sb)
require.NoError(t, err)
sb.Where(conds...)
sql, _ := sb.BuildWithFlavor(sqlbuilder.ClickHouse)
assert.Contains(t, sql, tc.expectedSQL)
assert.NotContains(t, sql, "scope.`scope.", "must not double-prefix the scope JSON path")
})
}
}
// TestConditionForSynthesizedKeys covers the KeyNotFound fallback: when a
// referenced attribute key has no metadata match, the builder synthesizes key(s) from
// user input and queries anyway, emitting a warning instead of failing.
@@ -504,17 +414,6 @@ func TestConditionForSynthesizedKeys(t *testing.T) {
assert.Contains(t, args, "timeout")
})
t.Run("scope context with no metadata -> scope attribute", func(t *testing.T) {
sb := sqlbuilder.NewSelectBuilder()
key := telemetrytypes.TelemetryFieldKey{Name: "custom.attr", FieldContext: telemetrytypes.FieldContextScope}
conds, warnings, err := cb.ConditionFor(ctx, valuer.UUID{}, 0, 0, &key, noMatches, qbtypes.ConditionBuilderOptions{}, qbtypes.FilterOperatorEqual, "v", sb)
assert.NoError(t, err, "an undeclared scope attribute must still be filterable")
assert.NotEmpty(t, warnings)
sb.Where(conds...)
sql, _ := sb.BuildWithFlavor(sqlbuilder.ClickHouse)
assert.Contains(t, sql, "scope.attributes.`custom.attr`")
})
t.Run("bare key with number operand -> attribute number", func(t *testing.T) {
sb := sqlbuilder.NewSelectBuilder()
key := telemetrytypes.TelemetryFieldKey{Name: "http.status"}

View File

@@ -121,20 +121,6 @@ var (
FieldContext: telemetrytypes.FieldContextSpan,
FieldDataType: telemetrytypes.FieldDataTypeString,
},
"scope.name": {
Name: "scope.name",
Description: "Instrumentation scope name",
Signal: telemetrytypes.SignalTraces,
FieldContext: telemetrytypes.FieldContextScope,
FieldDataType: telemetrytypes.FieldDataTypeString,
},
"scope.version": {
Name: "scope.version",
Description: "Instrumentation scope version",
Signal: telemetrytypes.SignalTraces,
FieldContext: telemetrytypes.FieldContextScope,
FieldDataType: telemetrytypes.FieldDataTypeString,
},
}
IntrinsicFieldsDeprecated = map[string]telemetrytypes.TelemetryFieldKey{
"traceID": {

View File

@@ -53,7 +53,6 @@ var (
ValueType: schema.ColumnTypeString,
}},
"resource": {Name: "resource", Type: schema.JSONColumnType{}},
"scope": {Name: "scope", Type: schema.JSONColumnType{}},
"events": {Name: "events", Type: schema.ArrayColumnType{
ElementType: schema.ColumnTypeString,
@@ -182,7 +181,7 @@ func (m *fieldMapper) getColumn(
case telemetrytypes.FieldContextResource:
return []*schema.Column{indexV3Columns["resource"], indexV3Columns["resources_string"]}, nil
case telemetrytypes.FieldContextScope:
return []*schema.Column{indexV3Columns["scope"]}, nil
return []*schema.Column{}, qbtypes.ErrColumnNotFound
case telemetrytypes.FieldContextAttribute:
switch key.FieldDataType {
case telemetrytypes.FieldDataTypeString:
@@ -293,25 +292,14 @@ func (m *fieldMapper) resolveColumnExprs(
switch column.Type.GetType() {
case schema.ColumnTypeEnumJSON:
// json is only supported for resource context as of now
if key.FieldContext != telemetrytypes.FieldContextResource {
return nil, nil, nil, errors.Newf(errors.TypeInvalidInput, errors.CodeInvalidInput, "only resource context fields are supported for json columns, got %s", key.FieldContext.String)
}
// have to add ::string as clickHouse throws an error :- data types Variant/Dynamic are not allowed in GROUP BY
// once clickHouse dependency is updated, we need to check if we can remove it.
switch key.FieldContext {
case telemetrytypes.FieldContextResource:
exprs = append(exprs, fmt.Sprintf("%s.`%s`::String", columnName, key.Name))
existExprs = append(existExprs, fmt.Sprintf("%s.`%s` IS NOT NULL", columnName, key.Name))
case telemetrytypes.FieldContextScope:
if f, ok := IntrinsicFields[key.Name]; ok && f.FieldContext == telemetrytypes.FieldContextScope {
// declared String paths on the scope column read '' for the missing case
exprs = append(exprs, fmt.Sprintf("%s::String", key.Name))
existExprs = append(existExprs, fmt.Sprintf("%s <> ''", key.Name))
} else {
attributeName := strings.TrimPrefix(key.Name, "attribute.") // literal "attribute" prefix in attribute keys needs double prefix
exprs = append(exprs, fmt.Sprintf("%s.attributes.%s::String", columnName, querybuilder.ClickHouseIdentifier(attributeName)))
existExprs = append(existExprs, fmt.Sprintf("%s.attributes.%s IS NOT NULL", columnName, querybuilder.ClickHouseIdentifier(attributeName)))
}
default:
return nil, nil, nil, errors.NewInternalf(errors.CodeInternal, "only resource and scope context fields are supported for json columns, got %s", key.FieldContext.String)
}
exprs = append(exprs, fmt.Sprintf("%s.`%s`::String", columnName, key.Name))
existExprs = append(existExprs, fmt.Sprintf("%s.`%s` IS NOT NULL", columnName, key.Name))
case schema.ColumnTypeEnumString,
schema.ColumnTypeEnumUInt64,
schema.ColumnTypeEnumUInt32,
@@ -353,6 +341,20 @@ func (m *fieldMapper) resolveColumnExprs(
return exprs, existExprs, columns, nil
}
// logicalForResolvedColumn upgrades a directly-resolvable key (the FieldFor
// probe succeeded) to its family when the metadata map proves membership;
// otherwise the key stays a single-member logical field.
func (m *fieldMapper) logicalForResolvedColumn(ctx context.Context, orgID valuer.UUID, field *telemetrytypes.TelemetryFieldKey, keys map[string][]*telemetrytypes.TelemetryFieldKey) *telemetrytypes.LogicalField {
for _, logical := range querybuilder.MatchingLogicalFields(ctx, orgID, m.fl, field, keys) {
if logical.IsFamily() &&
logical.FieldContext == field.FieldContext &&
(field.FieldDataType == telemetrytypes.FieldDataTypeUnspecified || logical.FieldDataType == field.FieldDataType) {
return logical
}
}
return telemetrytypes.SingleLogicalField(field.Name, field)
}
// upgradeToFamilies swaps single-member candidates for their family when the
// metadata map proves membership. Candidate order and every non-family
// candidate stay exactly as the legacy flow produced them; sibling candidates
@@ -417,12 +419,14 @@ func (m *fieldMapper) ColumnExpressionFor(
var candidates []*telemetrytypes.LogicalField
switch _, err := m.FieldFor(ctx, orgID, startNs, endNs, field); {
case err == nil:
// Every match from metadata is kept, similar to the filter path.
candidates = querybuilder.MatchingLogicalFields(ctx, orgID, m.fl, field, keys)
if len(candidates) == 0 {
candidates = []*telemetrytypes.LogicalField{telemetrytypes.SingleLogicalField(field.Name, field)}
}
// A directly-resolvable key upgrades to its family when the metadata
// map proves membership; otherwise it stays single-member.
candidates = []*telemetrytypes.LogicalField{m.logicalForResolvedColumn(ctx, orgID, field, keys)}
case errors.Is(err, qbtypes.ErrColumnNotFound):
// The legacy candidate flow, unchanged: column (when the bare name is
// one) plus metadata matches, else synthesized type-variant keys. The
// family step below only swaps candidates for their family; it never
// changes candidate order or non-family behavior.
raw := m.CandidateKeys(ctx, orgID, field, nil, keys)
if len(raw) == 0 {
return "", errors.Wrapf(err, errors.TypeInvalidInput, errors.CodeInvalidInput, "field `%s` not found", field.Name).WithSuggestions(errors.NewSuggestionsOnLevenshteinDistance(field.Name, errors.NounKeys, maps.Keys(keys))...)
@@ -595,37 +599,11 @@ func (m *fieldMapper) CandidateKeys(ctx context.Context, _ valuer.UUID, field *t
// strict context honored as-is: stripped interpretation first, literal spelling second
literal := telemetrytypes.NewTelemetryFieldKey(field.FieldContext.StringValue()+"."+field.Name, field.FieldContext, field.FieldDataType)
return append(querybuilder.SynthesizeKeys(field, value), querybuilder.SynthesizeKeys(literal, value)...)
case telemetrytypes.FieldContextScope:
// Declared scope paths (scope.name / scope.version) arrive via metadata as intrinsics;
// anything reaching synth is an undeclared scope attribute.
return []*telemetrytypes.TelemetryFieldKey{telemetrytypes.NewTelemetryFieldKey(field.Name, telemetrytypes.FieldContextScope, telemetrytypes.FieldDataTypeString)}
}
// contexts that don't exist on spans (log, body, …) have nothing to synthesize
// contexts that don't exist on spans (log, body, scope, …) have nothing to synthesize
return nil
}
// scopeJSONExistsExpression renders the existence predicate for the scope JSON column, the one
// signal-specific case the generic querybuilder.ExistsExpression must not carry.
func scopeJSONExistsExpression(key *telemetrytypes.TelemetryFieldKey, fieldExpression string, exists bool) (string, bool) {
if key.FieldContext != telemetrytypes.FieldContextScope {
return "", false
}
// Declared String paths are non-Nullable (absent reads '' not NULL).
if f, ok := IntrinsicFields[key.Name]; ok && f.FieldContext == telemetrytypes.FieldContextScope {
if exists {
return fieldExpression + " <> ''", true
}
return fieldExpression + " = ''", true
}
// Scope attribute: the value expression casts the JSON path to String, which folds a missing
// key's NULL to '', so presence must test the raw path — drop the ::String cast.
path := strings.TrimSuffix(fieldExpression, "::String")
if exists {
return path + " IS NOT NULL", true
}
return path + " IS NULL", true
}
// ExistsFor implements the per-key existence primitive of qbtypes.FieldMapper.
func (m *fieldMapper) ExistsFor(
ctx context.Context,
@@ -642,8 +620,5 @@ func (m *fieldMapper) ExistsFor(
if err != nil {
return "", err
}
if expr, ok := scopeJSONExistsExpression(key, fieldExpression, exists); ok {
return expr, nil
}
return querybuilder.ExistsExpression(columns, key, tsStart, tsEnd, fieldExpression, exists)
}

View File

@@ -84,45 +84,6 @@ func TestGetFieldKeyName(t *testing.T) {
expectedResult: "multiIf(resource.`deployment.environment` IS NOT NULL, resource.`deployment.environment`::String, `resource_string_deployment$$environment_exists`, `resource_string_deployment$$environment`, NULL)",
expectedError: nil,
},
{
name: "Scope field - scope.name",
key: telemetrytypes.TelemetryFieldKey{
Name: "scope.name",
FieldContext: telemetrytypes.FieldContextScope,
},
expectedResult: "scope.name::String",
expectedError: nil,
},
{
name: "Scope field - scope.version",
key: telemetrytypes.TelemetryFieldKey{
Name: "scope.version",
FieldContext: telemetrytypes.FieldContextScope,
},
expectedResult: "scope.version::String",
expectedError: nil,
},
{
name: "Scope field - custom attribute",
key: telemetrytypes.TelemetryFieldKey{
Name: "custom.attr",
FieldContext: telemetrytypes.FieldContextScope,
},
expectedResult: "scope.attributes.`custom.attr`::String",
expectedError: nil,
},
{
// `scope.attribute.name` normalizes to {attribute.name, scope}; the literal
// `attribute.` prefix is dropped so it addresses the scope attribute named `name`
// (which the declared `scope.name` path deliberately does not).
name: "Scope field - attribute prefix addresses the named scope attribute",
key: telemetrytypes.TelemetryFieldKey{
Name: "attribute.name",
FieldContext: telemetrytypes.FieldContextScope,
},
expectedResult: "scope.attributes.`name`::String",
expectedError: nil,
},
{
// Query like `attribute.attribute_string:string` should resolve to `attributes_string['attribute_string']`.
name: "Attribute key whose name collides with contextual map column resolves as a map lookup",
@@ -343,99 +304,3 @@ func TestColumnExpressionForTimestampAttributeCollision(t *testing.T) {
assert.Contains(t, result, "attributes_number['timestamp']")
})
}
// TestColumnExpressionForScopeDeclaredPath covers select-side resolution of scope names that
// collide with a declared scope path. A short name under scope context (or the bare
// `scope.<x>` spelling that normalizes to it) names both homes and coalesces them when
// metadata knows a same-named scope attribute, and binds to the declared path alone when it
// does not. The full `scope.<x>` name under explicit scope context addresses the declared
// path alone; the explicit `scope.attribute.` prefix addresses the attribute alone.
func TestColumnExpressionForScopeDeclaredPath(t *testing.T) {
ctx := context.Background()
scopeKey := func(name string) *telemetrytypes.TelemetryFieldKey {
return &telemetrytypes.TelemetryFieldKey{
Name: name,
Signal: telemetrytypes.SignalTraces,
FieldContext: telemetrytypes.FieldContextScope,
FieldDataType: telemetrytypes.FieldDataTypeString,
}
}
declaredOnly := map[string][]*telemetrytypes.TelemetryFieldKey{
"scope.name": {scopeKey("scope.name")},
"scope.version": {scopeKey("scope.version")},
}
withAttr := map[string][]*telemetrytypes.TelemetryFieldKey{
"scope.name": {scopeKey("scope.name")},
"scope.version": {scopeKey("scope.version")},
"name": {scopeKey("name")},
"version": {scopeKey("version")},
}
testCases := []struct {
name string
key telemetrytypes.TelemetryFieldKey
keys map[string][]*telemetrytypes.TelemetryFieldKey
expectedResult string
}{
{
name: "short name under scope context binds to the declared path",
key: telemetrytypes.TelemetryFieldKey{Name: "version", FieldContext: telemetrytypes.FieldContextScope},
keys: declaredOnly,
expectedResult: "multiIf(scope.version::String <> '', scope.version::String, NULL)",
},
{
name: "full scope.version name under scope context addresses the declared path alone",
key: telemetrytypes.TelemetryFieldKey{Name: "scope.version", FieldContext: telemetrytypes.FieldContextScope},
keys: withAttr,
expectedResult: "multiIf(scope.version::String <> '', scope.version::String, NULL)",
},
{
name: "full scope.name name under scope context addresses the declared path alone",
key: telemetrytypes.TelemetryFieldKey{Name: "scope.name", FieldContext: telemetrytypes.FieldContextScope},
keys: withAttr,
expectedResult: "multiIf(scope.name::String <> '', scope.name::String, NULL)",
},
{
// `scope.attribute.name` normalizes to {attribute.name, scope}; the `attribute.`
// prefix is dropped so it addresses the scope attribute named `name` — the only
// way to reach it, since `scope.name` is reserved for the declared path.
name: "attribute prefix reaches the named scope attribute",
key: telemetrytypes.TelemetryFieldKey{Name: "attribute.name", FieldContext: telemetrytypes.FieldContextScope},
keys: withAttr,
expectedResult: "multiIf(scope.attributes.`name` IS NOT NULL, scope.attributes.`name`::String, NULL)",
},
{
// the caller supplied the context, so `scope.` is part of the name rather than a
// prefix to strip: this addresses a scope attribute literally named
// `scope.testing.env`, not the attribute `testing.env`
name: "explicit context keeps a scope-prefixed name intact",
key: telemetrytypes.TelemetryFieldKey{Name: "scope.testing.env", FieldContext: telemetrytypes.FieldContextScope},
keys: declaredOnly,
expectedResult: "multiIf(scope.attributes.`scope.testing.env` IS NOT NULL, scope.attributes.`scope.testing.env`::String, NULL)",
},
{
// metadata knows both homes under this name, so the short spelling coalesces
// them instead of being rejected as ambiguous
name: "short name coalesces a known scope attribute with the declared path",
key: telemetrytypes.TelemetryFieldKey{Name: "name", FieldContext: telemetrytypes.FieldContextScope},
keys: withAttr,
expectedResult: "multiIf(scope.attributes.`name` IS NOT NULL, toString(scope.attributes.`name`::String), scope.name::String <> '', toString(scope.name::String), NULL)",
},
{
name: "short version coalesces a known scope attribute with the declared path",
key: telemetrytypes.TelemetryFieldKey{Name: "version", FieldContext: telemetrytypes.FieldContextScope},
keys: withAttr,
expectedResult: "multiIf(scope.attributes.`version` IS NOT NULL, toString(scope.attributes.`version`::String), scope.version::String <> '', toString(scope.version::String), NULL)",
},
}
for _, tc := range testCases {
t.Run(tc.name, func(t *testing.T) {
fm := NewFieldMapper(flaggertest.New(t))
result, err := fm.ColumnExpressionFor(ctx, valuer.UUID{}, 0, 0, &tc.key, telemetrytypes.FieldDataTypeUnspecified, tc.keys)
require.NoError(t, err)
assert.Equal(t, tc.expectedResult, result)
})
}
}

View File

@@ -113,20 +113,6 @@ func BuildCompleteFieldKeyMap(releaseTime time.Time) map[string][]*telemetrytype
FieldDataType: telemetrytypes.FieldDataTypeBool,
},
},
"scope.name": {
{
Name: "scope.name",
FieldContext: telemetrytypes.FieldContextScope,
FieldDataType: telemetrytypes.FieldDataTypeString,
},
},
"scope.version": {
{
Name: "scope.version",
FieldContext: telemetrytypes.FieldContextScope,
FieldDataType: telemetrytypes.FieldDataTypeString,
},
},
// both spellings of an enabled semantic-convention family
"deployment.environment.name": {
{

View File

@@ -18,7 +18,7 @@ import (
// - Use `scope.` prefix to explicitly indicate and enforce scope context. Example
// - `scope.name`
// - `scope.version`
// - `scope.my.custom.attribute` resolves to the `my.custom.attribute` scope attribute
// - `scope.my.custom.attribute` and `scope.attribute.my.custom.attribute` resolve to same attribute
//
// - Use `attribute.` to explicitly indicate and enforce attribute context. Example
// - `attribute.http.method`
@@ -190,7 +190,7 @@ func (FieldContext) Enum() []any {
FieldContextSpan,
FieldContextTrace,
FieldContextResource,
FieldContextScope,
// FieldContextScope,
FieldContextAttribute,
// FieldContextEvent,
FieldContextBody,

View File

@@ -294,17 +294,6 @@ func TestNormalize(t *testing.T) {
FieldDataType: FieldDataTypeString,
},
},
{
name: "Normalize keeps a prefix that does not match the set context",
input: TelemetryFieldKey{
Name: "scope.name",
FieldContext: FieldContextAttribute,
},
expected: TelemetryFieldKey{
Name: "scope.name",
FieldContext: FieldContextAttribute,
},
},
{
name: "Normalize body field",
input: TelemetryFieldKey{

View File

@@ -253,6 +253,27 @@ def get_preview_sql(response: requests.Response, name: str) -> str:
return statements[0]["db.statement.query"]
def get_preview_skip_indexes(response: requests.Response, name: str) -> dict[str, dict[str, Any]]:
"""The skip-index steps of the named query's read funnel, keyed by index name.
Needs a verbose preview. ClickHouse lists a skip index only when the predicate matches its
expression, so an absent entry means it was never consulted."""
statements = get_preview_statements(response, name)
assert len(statements) == 1, f"expected 1 statement for query {name}, got {len(statements)}"
granules = statements[0]["granules"]
assert granules is not None, f"query {name} reads no MergeTree table: {statements[0]}"
return {step["name"]: step for read in granules["reads"] for step in read["steps"] if step["type"] == "Skip"}
def get_preview_selected_granules(response: requests.Response, name: str) -> int:
"""Granules surviving every index step of the named query's read funnel."""
statements = get_preview_statements(response, name)
assert len(statements) == 1, f"expected 1 statement for query {name}, got {len(statements)}"
granules = statements[0]["granules"]
assert granules is not None, f"query {name} reads no MergeTree table: {statements[0]}"
return granules["selected"]
def aligned_epoch(ago: timedelta, step_seconds: int = DEFAULT_STEP_INTERVAL) -> int:
"""Epoch seconds for `now - ago`, floored to a step boundary so seeded
points land exactly on the query's toStartOfInterval buckets."""
@@ -999,8 +1020,6 @@ def generate_traces_with_corrupt_metadata() -> list[Traces]:
"cloud.provider": "integration",
"cloud.account.id": "000",
"trace_id": "corrupt_data",
"scope_name": "corrupt_data",
"scope.scope.name": "corrupt_data",
},
attributes={
"net.transport": "IP.TCP",
@@ -1009,10 +1028,7 @@ def generate_traces_with_corrupt_metadata() -> list[Traces]:
"http.request.method": "POST",
"http.response.status_code": "200",
"timestamp": "corrupt_data",
"version": "1.0.0",
"scope.scope.version": "1.0.0",
},
scope={"name": "io.signoz.http.server", "version": "2.0.0"},
),
Traces(
timestamp=now - timedelta(seconds=3.5),
@@ -1032,24 +1048,12 @@ def generate_traces_with_corrupt_metadata() -> list[Traces]:
"cloud.provider": "integration",
"cloud.account.id": "000",
"timestamp": "corrupt_data",
"scope.attributes.name": "corrupt_data",
},
attributes={
"db.name": "integration",
"db.operation": "SELECT",
"db.statement": "SELECT * FROM integration",
"trace_d": "corrupt_data",
"scope.attributes.version": "corrupt_data",
},
scope={
"name": "io.opentelemetry.contrib.http",
"version": "1.0.0",
"attributes": {
"telemetry.sdk.language": "cpp",
"name": "not-the-real-name",
"version": "not-the-real-version",
"attributes": "literally-a-key-named-attributes",
},
},
),
Traces(
@@ -1070,15 +1074,12 @@ def generate_traces_with_corrupt_metadata() -> list[Traces]:
"cloud.provider": "integration",
"cloud.account.id": "000",
"duration_nano": "corrupt_data",
"scope.scope.attributes.version": "corrupt_data",
},
attributes={
"http.request.method": "PATCH",
"http.status_code": "404",
"id": "1",
"scope.scope.version": "corrupt_data",
},
scope={"name": "io.signoz.http.client", "version": "2.0.0"},
),
Traces(
timestamp=now - timedelta(seconds=1),
@@ -1097,7 +1098,6 @@ def generate_traces_with_corrupt_metadata() -> list[Traces]:
"host.name": "linux-001",
"cloud.provider": "integration",
"cloud.account.id": "001",
"scope.scope.version": "corrupt_data",
},
attributes={
"message.type": "SENT",
@@ -1105,10 +1105,7 @@ def generate_traces_with_corrupt_metadata() -> list[Traces]:
"messaging.message.id": "001",
"duration_nano": "corrupt_data",
"id": 1,
"scope": "corrupt_data",
"scope.attributes.name": "corrupt_data",
},
scope={"name": "io.signoz.messaging", "version": "3.0.0"},
),
]

View File

@@ -302,7 +302,6 @@ class Traces(ABC):
db_operation: str
has_error: bool
is_remote: str
scope_json: dict[str, Any]
resource: list[TracesResource]
tag_attributes: list[TracesTagAttributes]
@@ -328,7 +327,6 @@ class Traces(ABC):
links: list[TracesLink] = [],
trace_state: str = "",
flags: np.uint32 = 0,
scope: dict[str, Any] = {},
resource_write_mode: Literal["legacy_only", "dual_write"] = "dual_write",
) -> None:
if timestamp is None:
@@ -410,33 +408,6 @@ class Traces(ABC):
# Calculate resource fingerprint
self.resource_fingerprint = LogsOrTracesFingerprint(self.resources_string).calculate()
# Process scope mirroring the InstrumentationScope on the OTLP span.
scope_name = scope.get("name", "")
scope_version = scope.get("version", "")
scope_string = {k: str(v) for k, v in scope.get("attributes", {}).items()}
self.scope_json = {
"name": scope_name,
"version": scope_version,
"attributes": scope_string,
}
scope_keys = {"scope.name": scope_name, "scope.version": scope_version}
scope_keys.update(scope_string)
for k, v in scope_keys.items():
if v == "":
continue
self.tag_attributes.append(
TracesTagAttributes(
timestamp=timestamp,
tag_key=k,
tag_type="scope",
tag_data_type="string",
string_value=v,
number_value=None,
)
)
self.attribute_keys.append(TracesResourceOrAttributeKeys(name=k, datatype="string", tag_type="scope"))
# Process attributes by type and populate custom fields
self.attribute_string = {}
self.attributes_number = {}
@@ -688,7 +659,6 @@ class Traces(ABC):
self.has_error,
self.is_remote,
self.resource_json,
self.scope_json,
],
dtype=object,
)
@@ -719,7 +689,6 @@ class Traces(ABC):
attributes=data.get("attributes", {}),
trace_state=data.get("trace_state", ""),
flags=data.get("flags", 0),
scope=data.get("scope", {}),
)
@classmethod
@@ -859,7 +828,6 @@ def insert_traces_to_clickhouse(conn, traces: list[Traces]) -> None:
"has_error",
"is_remote",
"resource",
"scope",
],
data=[trace.np_arr() for trace in traces],
)

View File

@@ -0,0 +1,90 @@
from collections.abc import Callable
from datetime import UTC, datetime, timedelta
from http import HTTPStatus
import pytest
from fixtures import types
from fixtures.auth import USER_ADMIN_EMAIL, USER_ADMIN_PASSWORD
from fixtures.logs import Logs
from fixtures.querier import build_order_by, build_raw_query, get_rows, make_query_request
LOWER = "alpha"
UPPER = "ALPHA"
PLAIN = "beta"
NON_ASCII = "Mixed CASE Ünïcode"
SLASH = "GET /api/v1/users"
SUPERSTRING = "GET /api/v1/users/42"
QUOTE = 'say "hi" now'
BACKSLASH = "C:\\tmp\\log"
LIKE_META = "100% _off"
TAB = "tab\there"
CTRL = "ctrl\x01here"
BODIES = [LOWER, UPPER, PLAIN, NON_ASCII, SLASH, SUPERSTRING, QUOTE, BACKSLASH, LIKE_META, TAB, CTRL]
# querierlogs/16_body_equality.py with use_json_body on: `body` resolves to body_v2.message,
# which the lower(body) companion skips, and the same expressions must still answer alike.
@pytest.mark.parametrize(
"expression,expected_bodies",
[
pytest.param(f"body = '{LOWER}'", {LOWER}, id="equality_exact"),
pytest.param(f"body = '{UPPER}'", {UPPER}, id="equality_other_case"),
pytest.param("body = 'Alpha'", set(), id="equality_case_must_match"),
pytest.param(f"body = '{NON_ASCII}'", {NON_ASCII}, id="equality_non_ascii"),
pytest.param("body = 'gamma'", set(), id="equality_no_match"),
pytest.param(f"body = '{SLASH}'", {SLASH}, id="equality_slash"),
pytest.param("body = 'say \"hi\" now'", {QUOTE}, id="equality_quote"),
pytest.param(r"body = 'C:\\tmp\\log'", {BACKSLASH}, id="equality_backslash"),
pytest.param(f"body = '{LIKE_META}'", {LIKE_META}, id="equality_like_metacharacters"),
pytest.param("body = 'tab\there'", {TAB}, id="equality_tab"),
pytest.param("body = 'ctrl\x01here'", {CTRL}, id="equality_control_char"),
pytest.param("body = 'GET /api/v1'", set(), id="equality_prefix_does_not_match"),
pytest.param(f"body IN ('{LOWER}', '{PLAIN}')", {LOWER, PLAIN}, id="in_excludes_other_case"),
pytest.param(f"body IN ('{SLASH}', '{LIKE_META}')", {SLASH, LIKE_META}, id="in_escaped_values"),
pytest.param(f"body NOT IN ('{LOWER}', '{UPPER}')", set(BODIES) - {LOWER, UPPER}, id="not_in"),
],
)
def test_logs_body_equality_json(
signoz: types.SigNoz,
create_user_admin: None, # pylint: disable=unused-argument
get_token: Callable[[str, str], str],
insert_logs: Callable[[list[Logs]], None],
expression: str,
expected_bodies: set[str],
) -> None:
now = datetime.now(tz=UTC).replace(second=0, microsecond=0)
insert_logs(
[
Logs(
timestamp=now - timedelta(seconds=i + 1),
resources={"service.name": "api"},
body=body,
)
for i, body in enumerate(BODIES)
]
)
token = get_token(USER_ADMIN_EMAIL, USER_ADMIN_PASSWORD)
response = make_query_request(
signoz,
token,
start_ms=int((now - timedelta(minutes=5)).timestamp() * 1000),
end_ms=int(now.timestamp() * 1000),
request_type="raw",
queries=[
build_raw_query(
"A",
"logs",
filter_expression=expression,
order=[build_order_by("timestamp", "desc"), build_order_by("id", "desc")],
limit=100,
)
],
)
assert response.status_code == HTTPStatus.OK, response.text
assert response.json()["status"] == "success"
# body_v2 comes back parsed; a plain-string body is {"message": <body>}.
assert {row["data"]["body"]["message"] for row in get_rows(response)} == expected_bodies

View File

@@ -0,0 +1,92 @@
from collections.abc import Callable
from datetime import UTC, datetime, timedelta
from http import HTTPStatus
import pytest
from fixtures import types
from fixtures.auth import USER_ADMIN_EMAIL, USER_ADMIN_PASSWORD
from fixtures.logs import Logs
from fixtures.querier import build_order_by, build_raw_query, get_column_data_from_response, make_query_request
LOWER = "alpha"
UPPER = "ALPHA"
PLAIN = "beta"
NON_ASCII = "Mixed CASE Ünïcode"
SLASH = "GET /api/v1/users"
SUPERSTRING = "GET /api/v1/users/42"
QUOTE = 'say "hi" now'
BACKSLASH = "C:\\tmp\\log"
LIKE_META = "100% _off"
TAB = "tab\there"
CTRL = "ctrl\x01here"
BODIES = [LOWER, UPPER, PLAIN, NON_ASCII, SLASH, SUPERSTRING, QUOTE, BACKSLASH, LIKE_META, TAB, CTRL]
# `body = ?` carries a case-insensitive LOWER(body) companion for the bloom filters, so a
# body differing only in case must still not come back.
@pytest.mark.parametrize(
"expression,expected_bodies",
[
pytest.param(f"body = '{LOWER}'", {LOWER}, id="equality_exact"),
pytest.param(f"body = '{UPPER}'", {UPPER}, id="equality_other_case"),
pytest.param("body = 'Alpha'", set(), id="equality_case_must_match"),
pytest.param(f"body = '{NON_ASCII}'", {NON_ASCII}, id="equality_non_ascii"),
pytest.param("body = ''", set(), id="equality_empty"),
pytest.param("body = 'gamma'", set(), id="equality_no_match"),
# the companion is a LIKE-free equality, so none of these are metacharacters to it
pytest.param(f"body = '{SLASH}'", {SLASH}, id="equality_slash"),
pytest.param("body = 'say \"hi\" now'", {QUOTE}, id="equality_quote"),
pytest.param(r"body = 'C:\\tmp\\log'", {BACKSLASH}, id="equality_backslash"),
pytest.param(f"body = '{LIKE_META}'", {LIKE_META}, id="equality_like_metacharacters"),
pytest.param("body = 'tab\there'", {TAB}, id="equality_tab"),
pytest.param("body = 'ctrl\x01here'", {CTRL}, id="equality_control_char"),
# a prefix of another body must not match it
pytest.param("body = 'GET /api/v1'", set(), id="equality_prefix_does_not_match"),
pytest.param(f"body IN ('{LOWER}', '{PLAIN}')", {LOWER, PLAIN}, id="in_excludes_other_case"),
pytest.param(f"body IN ('{SLASH}', '{LIKE_META}')", {SLASH, LIKE_META}, id="in_escaped_values"),
pytest.param(f"body NOT IN ('{LOWER}', '{UPPER}')", set(BODIES) - {LOWER, UPPER}, id="not_in"),
],
)
def test_logs_body_equality(
signoz: types.SigNoz,
create_user_admin: None, # pylint: disable=unused-argument
get_token: Callable[[str, str], str],
insert_logs: Callable[[list[Logs]], None],
expression: str,
expected_bodies: set[str],
) -> None:
now = datetime.now(tz=UTC).replace(second=0, microsecond=0)
insert_logs(
[
Logs(
timestamp=now - timedelta(seconds=i + 1),
resources={"service.name": "api"},
body=body,
)
for i, body in enumerate(BODIES)
]
)
token = get_token(USER_ADMIN_EMAIL, USER_ADMIN_PASSWORD)
response = make_query_request(
signoz,
token,
start_ms=int((now - timedelta(minutes=5)).timestamp() * 1000),
end_ms=int(now.timestamp() * 1000),
request_type="raw",
queries=[
build_raw_query(
"A",
"logs",
filter_expression=expression,
order=[build_order_by("timestamp", "desc"), build_order_by("id", "desc")],
limit=100,
)
],
)
assert response.status_code == HTTPStatus.OK, response.text
assert response.json()["status"] == "success"
assert set(get_column_data_from_response(response.json(), "body")) == expected_bodies

View File

@@ -0,0 +1,346 @@
import json
from collections.abc import Callable
from datetime import UTC, datetime, timedelta
from http import HTTPStatus
import pytest
from fixtures import types
from fixtures.auth import USER_ADMIN_EMAIL, USER_ADMIN_PASSWORD
from fixtures.logs import Logs
from fixtures.querier import (
build_order_by,
build_raw_query,
get_preview_selected_granules,
get_preview_skip_indexes,
get_rows,
make_preview_query_request,
make_query_request,
)
# The legacy body holds the text the producer wrote, so the same value reaches ClickHouse under
# different encodings: PHP escapes `/`, Go escapes `&` `<` `>`, Python escapes non-ASCII. The
# LOWER(body) predicates the filters carry for the bloom filters must find all of them.
BODIES = {
"plain": '{"tag":"plain","url":"https://signoz.io/docs","user_id":4242,"status":"timeout_error"}',
"php": '{"tag":"php","url":"https:\\/\\/signoz.io\\/docs"}',
"go": '{"tag":"go","note":"connection reset \\u0026 retry aborted"}',
"python": '{"tag":"python","city":"caf\\u00e9 municipal district"}',
"other_case": '{"tag":"other_case","status":"TIMEOUT_ERROR"}',
"no_user_id": '{"tag":"no_user_id","status":"ok","url":"https://signoz.io/pricing"}',
"tagged": '{"tag":"tagged","labels":["production","webserver"]}',
"tagged_escaped": '{"tag":"tagged_escaped","labels":["batch \\u0026 stream","webserver"]}',
}
@pytest.mark.parametrize(
"expression,expected_tags",
[
pytest.param("body.user_id = 4242", {"plain"}, id="numeric_equality"),
pytest.param("body.user_id EXISTS", {"plain"}, id="exists"),
# a negated comparison matches the rows without the path, so it carries no predicate
pytest.param("body.user_id != 4242", set(BODIES) - {"plain"}, id="not_equal_keeps_pathless_rows"),
pytest.param("body.status NOT EXISTS", {"php", "go", "python", "tagged", "tagged_escaped"}, id="not_exists"),
pytest.param("body.url = 'https://signoz.io/docs'", {"plain", "php"}, id="equality_escaped_slashes"),
pytest.param("body.url CONTAINS 'signoz.io/docs'", {"plain", "php"}, id="contains_escaped_slashes"),
pytest.param("body.note = 'connection reset & retry aborted'", {"go"}, id="equality_escaped_ampersand"),
pytest.param("body.city = 'café municipal district'", {"python"}, id="equality_escaped_non_ascii"),
# the value predicate is case-insensitive where the equality is not
pytest.param("body.status = 'timeout_error'", {"plain"}, id="equality_underscore"),
pytest.param("body.status = 'TIMEOUT_ERROR'", {"other_case"}, id="equality_other_case"),
pytest.param("body.status IN ('timeout_error', 'ok')", {"plain", "no_user_id"}, id="in_carries_one_value_per_arm"),
# has and hasAll assert every element, hasAny only one of them
pytest.param("has(body.labels[*], 'production')", {"tagged"}, id="has_element"),
pytest.param("has(body.labels[*], 'batch & stream')", {"tagged_escaped"}, id="has_escaped_element"),
pytest.param("hasAll(body.labels[*], ['production', 'webserver'])", {"tagged"}, id="has_all_needs_every_element"),
pytest.param(
"hasAny(body.labels[*], ['production', 'batch & stream'])",
{"tagged", "tagged_escaped"},
id="has_any_needs_one_element",
),
# 'webserver' is in both, so an ORed literal set must not exclude either row
pytest.param(
"hasAny(body.labels[*], ['webserver', 'nothing here'])",
{"tagged", "tagged_escaped"},
id="has_any_across_both",
),
],
)
def test_logs_body_json_index_predicates(
signoz: types.SigNoz,
create_user_admin: None, # pylint: disable=unused-argument
get_token: Callable[[str, str], str],
insert_logs: Callable[[list[Logs]], None],
expression: str,
expected_tags: set[str],
) -> None:
now = datetime.now(tz=UTC).replace(second=0, microsecond=0)
insert_logs(
[
Logs(
timestamp=now - timedelta(seconds=i + 1),
resources={"service.name": "api"},
body=body,
)
for i, body in enumerate(BODIES.values())
]
)
token = get_token(USER_ADMIN_EMAIL, USER_ADMIN_PASSWORD)
response = make_query_request(
signoz,
token,
start_ms=int((now - timedelta(minutes=5)).timestamp() * 1000),
end_ms=int(now.timestamp() * 1000),
request_type="raw",
queries=[
build_raw_query(
"A",
"logs",
filter_expression=expression,
order=[build_order_by("timestamp", "desc"), build_order_by("id", "desc")],
limit=100,
)
],
)
assert response.status_code == HTTPStatus.OK, response.text
assert response.json()["status"] == "success"
assert {json.loads(row["data"]["body"])["tag"] for row in get_rows(response)} == expected_tags
# JSON_VALUE matches no index expression, so the literals are what get the bloom filters consulted
# at all; the read funnel is what catches one that stops matching.
BODY_BLOOM_FILTERS = {"body_index_v2_token", "body_index_v2_ngram"}
@pytest.mark.parametrize(
"expression,prunes_every_granule",
[
pytest.param("body.status = 'timeout_error'", False, id="value_needle_present"),
pytest.param("body.status = 'zz_no_seeded_body_holds_this'", True, id="value_needle_absent"),
pytest.param("body.zz_no_seeded_body_holds_this EXISTS", True, id="path_needle_absent"),
# a number is compared after JSONExtract parses it, so it carries no value literal - only
# its path, which every row holding the key satisfies
pytest.param("body.user_id = 999999", False, id="number_carries_only_its_path"),
],
)
def test_logs_body_json_index_prunes_granules(
signoz: types.SigNoz,
create_user_admin: None, # pylint: disable=unused-argument
get_token: Callable[[str, str], str],
insert_logs: Callable[[list[Logs]], None],
expression: str,
prunes_every_granule: bool,
) -> None:
now = datetime.now(tz=UTC).replace(second=0, microsecond=0)
insert_logs(
[
Logs(
timestamp=now - timedelta(seconds=i + 1),
resources={"service.name": "api"},
body=body,
)
for i, body in enumerate(BODIES.values())
]
)
token = get_token(USER_ADMIN_EMAIL, USER_ADMIN_PASSWORD)
response = make_preview_query_request(
signoz,
token,
start_ms=int((now - timedelta(minutes=5)).timestamp() * 1000),
end_ms=int(now.timestamp() * 1000),
request_type="raw",
queries=[
build_raw_query(
"A",
"logs",
filter_expression=expression,
order=[build_order_by("timestamp", "desc"), build_order_by("id", "desc")],
limit=100,
)
],
)
assert response.status_code == HTTPStatus.OK, response.text
skip_indexes = get_preview_skip_indexes(response, "A")
assert BODY_BLOOM_FILTERS <= set(skip_indexes), f"body bloom filters not consulted, only: {sorted(skip_indexes)}"
selected = get_preview_selected_granules(response, "A")
if prunes_every_granule:
assert selected == 0, f"expected every granule pruned, {selected} survived"
else:
assert selected > 0, "the granule holding the match must survive"
# `body = ?` matches no index expression on its own; the lowered companion is what the filters
# prune on, so its absence from the funnel is the regression this catches.
@pytest.mark.parametrize(
"expression,prunes_every_granule",
[
pytest.param("body = 'alpha'", False, id="equality_present_value"),
pytest.param("body = 'zz_no_seeded_body_holds_this'", True, id="equality_absent_value"),
# IN delegates to the equalities, so every arm carries its own companion
pytest.param("body IN ('alpha', 'beta')", False, id="in_present_values"),
pytest.param("body IN ('zz_absent_one', 'zz_absent_two')", True, id="in_absent_values"),
],
)
def test_logs_body_equality_prunes_granules(
signoz: types.SigNoz,
create_user_admin: None, # pylint: disable=unused-argument
get_token: Callable[[str, str], str],
insert_logs: Callable[[list[Logs]], None],
expression: str,
prunes_every_granule: bool,
) -> None:
now = datetime.now(tz=UTC).replace(second=0, microsecond=0)
insert_logs([Logs(timestamp=now - timedelta(seconds=i + 1), resources={"service.name": "api"}, body=body) for i, body in enumerate(["alpha", "ALPHA", "beta"])])
token = get_token(USER_ADMIN_EMAIL, USER_ADMIN_PASSWORD)
response = make_preview_query_request(
signoz,
token,
start_ms=int((now - timedelta(minutes=5)).timestamp() * 1000),
end_ms=int(now.timestamp() * 1000),
request_type="raw",
queries=[
build_raw_query(
"A",
"logs",
filter_expression=expression,
order=[build_order_by("timestamp", "desc"), build_order_by("id", "desc")],
limit=100,
)
],
)
assert response.status_code == HTTPStatus.OK, response.text
skip_indexes = get_preview_skip_indexes(response, "A")
assert BODY_BLOOM_FILTERS <= set(skip_indexes), f"body bloom filters not consulted, only: {sorted(skip_indexes)}"
selected = get_preview_selected_granules(response, "A")
if prunes_every_granule:
assert selected == 0, f"expected every granule pruned, {selected} survived"
else:
assert selected > 0, "the granule holding the match must survive"
# The attribute maps carry a bloom filter over mapValues, which the subscript the comparison reads
# matches no more than the body column did. mapContains prunes on the key alone, so these use a key
# every row carries to isolate what the value predicate contributes.
ATTRIBUTE_NUMBER_VALUE_INDEX = "attributes_number_idx_val"
ATTRIBUTE_STRING_VALUE_INDEX = "attributes_string_idx_val"
@pytest.mark.parametrize(
"expression,prunes_every_granule",
[
pytest.param("attribute.resp_code = 503", False, id="value_present"),
pytest.param("attribute.resp_code = 60599", True, id="value_absent"),
# IN delegates to the equalities, so every arm carries its own membership assertion
pytest.param("attribute.resp_code IN (503, 200)", False, id="in_present_values"),
pytest.param("attribute.resp_code IN (60599, 60600)", True, id="in_absent_values"),
],
)
def test_attribute_number_equality_prunes_granules(
signoz: types.SigNoz,
create_user_admin: None, # pylint: disable=unused-argument
get_token: Callable[[str, str], str],
insert_logs: Callable[[list[Logs]], None],
expression: str,
prunes_every_granule: bool,
) -> None:
now = datetime.now(tz=UTC).replace(second=0, microsecond=0)
insert_logs([Logs(timestamp=now - timedelta(seconds=i + 1), resources={"service.name": "api"}, attributes={"resp_code": code}) for i, code in enumerate([200, 200, 503])])
token = get_token(USER_ADMIN_EMAIL, USER_ADMIN_PASSWORD)
response = make_preview_query_request(
signoz,
token,
start_ms=int((now - timedelta(minutes=5)).timestamp() * 1000),
end_ms=int(now.timestamp() * 1000),
request_type="raw",
queries=[
build_raw_query(
"A",
"logs",
filter_expression=expression,
order=[build_order_by("timestamp", "desc"), build_order_by("id", "desc")],
limit=100,
)
],
)
assert response.status_code == HTTPStatus.OK, response.text
skip_indexes = get_preview_skip_indexes(response, "A")
assert ATTRIBUTE_NUMBER_VALUE_INDEX in skip_indexes, f"mapValues filter not consulted, only: {sorted(skip_indexes)}"
selected = get_preview_selected_granules(response, "A")
if prunes_every_granule:
assert selected == 0, f"expected every granule pruned, {selected} survived"
else:
assert selected > 0, "the granule holding the match must survive"
# The mapValues filter indexes the values raw, so a case-insensitive match reaches it only for a
# pattern holding no ASCII letter, where LOWER changes nothing. A letter leaves it unconsulted.
@pytest.mark.parametrize(
"expression,value_index_consulted,prunes_every_granule",
[
pytest.param("attribute.client.ip CONTAINS '192.168.77'", True, False, id="letter_free_value_present"),
pytest.param("attribute.client.ip CONTAINS '192.168.99'", True, True, id="letter_free_value_absent"),
pytest.param("attribute.env CONTAINS 'production'", False, False, id="letters_leave_it_unconsulted"),
# the filter is consulted for any letter-free pattern, but a run below the index ngram
# leaves it nothing to check
pytest.param("attribute.client.ip CONTAINS '.7'", True, False, id="run_shorter_than_the_ngram"),
],
)
def test_attribute_letter_free_match_prunes_granules(
signoz: types.SigNoz,
create_user_admin: None, # pylint: disable=unused-argument
get_token: Callable[[str, str], str],
insert_logs: Callable[[list[Logs]], None],
expression: str,
value_index_consulted: bool,
prunes_every_granule: bool,
) -> None:
now = datetime.now(tz=UTC).replace(second=0, microsecond=0)
insert_logs(
[
Logs(
timestamp=now - timedelta(seconds=i + 1),
resources={"service.name": "api"},
attributes={"client.ip": ip, "env": "production"},
)
for i, ip in enumerate(["10.0.0.1", "10.0.0.2", "192.168.77.31"])
]
)
token = get_token(USER_ADMIN_EMAIL, USER_ADMIN_PASSWORD)
response = make_preview_query_request(
signoz,
token,
start_ms=int((now - timedelta(minutes=5)).timestamp() * 1000),
end_ms=int(now.timestamp() * 1000),
request_type="raw",
queries=[
build_raw_query(
"A",
"logs",
filter_expression=expression,
order=[build_order_by("timestamp", "desc"), build_order_by("id", "desc")],
limit=100,
)
],
)
assert response.status_code == HTTPStatus.OK, response.text
skip_indexes = get_preview_skip_indexes(response, "A")
assert (ATTRIBUTE_STRING_VALUE_INDEX in skip_indexes) == value_index_consulted, f"consulted: {sorted(skip_indexes)}"
selected = get_preview_selected_granules(response, "A")
if prunes_every_granule:
assert selected == 0, f"expected every granule pruned, {selected} survived"
else:
assert selected > 0, "the granule holding the match must survive"

View File

@@ -1240,13 +1240,6 @@ def test_traces_list_span_scope(
lambda x: {"duration_nano": int(x[1].duration_nano), "span_id": x[1].span_id, "timestamp": format_timestamp(x[1].timestamp), "trace_id": x[1].trace_id},
id="select_attribute_duration_order_intrinsic",
),
# Case 9: filter on the intrinsic scope.version. Only x[1] should match.
pytest.param(
BuilderQuery(signal="traces", name="A", select_fields=[TelemetryFieldKey("timestamp")], filter_expression="scope.version = '1.0.0'", limit=1),
HTTPStatus.OK,
lambda x: {"span_id": x[1].span_id, "timestamp": format_timestamp(x[1].timestamp), "trace_id": x[1].trace_id},
id="filter_scope_version",
),
],
)
def test_traces_list_with_corrupt_data(
@@ -1290,168 +1283,6 @@ def test_traces_list_with_corrupt_data(
assert get_rows(response)[0]["data"] == expected(traces)
@pytest.mark.parametrize(
"filter_expression,expected_indices",
[
# Intrinsic scope.name / scope.version resolve to the JSON sub-columns.
pytest.param("scope.name = 'io.signoz.payment'", [1], id="intrinsic_scope_name"),
pytest.param("scope.version = '2.3.1'", [0], id="intrinsic_scope_version"),
# A scope attribute resolves against the scope JSON column's attributes.
pytest.param("scope.telemetry.sdk.language = 'python'", [1], id="scope_attribute"),
# A scope attribute whose own name carries a `scope.` prefix. `scope.prefixed`
# normalizes to {prefixed, scope} and must still resolve to the attribute.
pytest.param("scope.prefixed = 'prefixed-val'", [0], id="scope_prefixed_attribute"),
# `env.tier` is a span attribute on span 0 and a scope attribute on
# span 1. Unprefixed -> no explicit context, so it is checked in every
# applicable context (attribute OR scope) and both spans match.
pytest.param("env.tier = 'gold'", [0, 1], id="bare_cross_context"),
# The explicit `scope.` prefix forces scope context only, so span 0's
# span attribute is ignored — only span 1 matches.
pytest.param("scope.env.tier = 'gold'", [1], id="scope_prefixed_cross_context"),
# `scope.name` names both homes it can resolve to: the declared scope.name field
# (span 0) and a same-named `name` scope attribute (span 1), the same way any
# other name colliding across contexts unions. `scope.attribute.name` addresses
# the attribute alone.
pytest.param("scope.name = 'io.signoz.checkout'", [0, 1], id="scope_name_unions_attribute"),
# The `scope.name` spelling is also a real stored key: span 2 carries a span
# attribute literally named `scope.name`, so it matches too.
pytest.param("scope.name = 'attr-scope-name'", [2], id="scope_name_matches_stored_spelling"),
# The explicit `scope.attribute.` prefix addresses the scope attribute alone, without
# the declared path. Span 1 has a `name` scope attribute = 'io.signoz.checkout'.
pytest.param("scope.attribute.name = 'io.signoz.checkout'", [1], id="scope_attribute_name"),
# `version` as a scope attribute: no span carries one (span 1's 4.5.6 is the declared
# scope.version, not a scope attribute), so this matches nothing.
pytest.param("scope.attribute.version = '4.5.6'", [], id="scope_attribute_version_none"),
# An unprefixed `name` is checked in every applicable context: the span `name`
# column (span 2) and a `name` scope attribute (span 1). It does not reach the
# declared scope.name field (span 0), which only the `scope.` prefix addresses.
pytest.param("name = 'io.signoz.checkout'", [1, 2], id="bare_name_unions_scope_attribute"),
# A value that no resolvable key holds (scope.name/scope.version field,
# a `name`/`version` scope attribute, or a same-named attribute/resource)
# returns nothing.
pytest.param("scope.version = 'corrupt_data'", [], id="scope_version_no_match"),
pytest.param("scope.name = 'corrupt_data'", [], id="scope_name_no_match"),
],
)
def test_traces_list_with_scope_filter(
signoz: types.SigNoz,
create_user_admin: None, # pylint: disable=unused-argument
get_token: Callable[[str, str], str],
insert_traces: Callable[[list[Traces]], None],
filter_expression: str,
expected_indices: list[int],
) -> None:
"""
Setup three spans with different scope key resolution:
- x[0]: scope.name/version 'io.signoz.checkout'/'2.3.1'; span attribute
env.tier='gold'.
- x[1]: scope.name/version 'io.signoz.payment'/'4.5.6'; scope attributes
telemetry.sdk.language='python', env.tier='gold', and a `name` scope
attribute colliding with x[0]'s scope.name value.
- x[2]: span name 'io.signoz.checkout' (colliding with x[0]'s scope.name
value) and a span attribute literally named `scope.name`.
Tests:
- Filtering on scope.name / scope.version / a scope attribute.
- An unprefixed key is resolved across contexts (scope checked alongside
attribute / intrinsic), while a `scope.`-prefixed key is scope-only.
- `scope.name`/`scope.version` name every home they resolve to: the declared JSON
sub-column and a same-named `name`/`version` scope attribute. The explicit
`scope.attribute.` prefix addresses the attribute alone.
- a bare `name` reaches the span `name` column and a `name` scope attribute, but
never the declared scope.name field.
"""
now = datetime.now(tz=UTC).replace(microsecond=0)
trace_id = TraceIdGenerator.trace_id()
span_ids = [TraceIdGenerator.span_id() for _ in range(3)]
traces = [
Traces(
timestamp=now - timedelta(seconds=4),
duration=timedelta(seconds=2),
trace_id=trace_id,
span_id=span_ids[0],
parent_span_id="",
name="GET /checkout",
kind=TracesKind.SPAN_KIND_SERVER,
status_code=TracesStatusCode.STATUS_CODE_OK,
resources={"service.name": "checkout"},
attributes={"http.request.method": "GET", "env.tier": "gold"},
scope={
"name": "io.signoz.checkout",
"version": "2.3.1",
# a scope attribute whose own name carries a `scope.` prefix
"attributes": {"telemetry.sdk.language": "go", "scope.prefixed": "prefixed-val"},
},
),
Traces(
timestamp=now - timedelta(seconds=2),
duration=timedelta(seconds=1),
trace_id=trace_id,
span_id=span_ids[1],
parent_span_id="",
name="POST /pay",
kind=TracesKind.SPAN_KIND_SERVER,
status_code=TracesStatusCode.STATUS_CODE_OK,
resources={"service.name": "payment"},
attributes={"http.request.method": "POST"},
# env.tier is a scope attribute here (cross-context with span 0);
# `name` is a scope attribute colliding with span 0's scope.name.
scope={
"name": "io.signoz.payment",
"version": "4.5.6",
"attributes": {
"telemetry.sdk.language": "python",
"env.tier": "gold",
"name": "io.signoz.checkout",
},
},
),
Traces(
timestamp=now - timedelta(seconds=1),
duration=timedelta(seconds=1),
trace_id=trace_id,
span_id=span_ids[2],
parent_span_id="",
# span name collides with span 0's scope.name value
name="io.signoz.checkout",
kind=TracesKind.SPAN_KIND_SERVER,
status_code=TracesStatusCode.STATUS_CODE_OK,
resources={"service.name": "probe"},
# a span attribute named `scope.name`
attributes={"scope.name": "attr-scope-name"},
scope={"name": "span-gamma", "version": "9.9.9"},
),
]
insert_traces(traces)
token = get_token(USER_ADMIN_EMAIL, USER_ADMIN_PASSWORD)
start_ms = int((now - timedelta(minutes=1)).timestamp() * 1000)
end_ms = int((now + timedelta(seconds=1)).timestamp() * 1000)
response = make_query_request(
signoz,
token,
start_ms=start_ms,
end_ms=end_ms,
request_type=RequestType.RAW,
queries=[
BuilderQuery(
signal="traces",
name="A",
select_fields=[TelemetryFieldKey("timestamp")],
filter_expression=filter_expression,
limit=10,
).to_dict()
],
)
assert response.status_code == HTTPStatus.OK, response.text
got_span_ids = {row["data"]["span_id"] for row in get_rows(response)}
expected_span_ids = {traces[i].span_id for i in expected_indices}
assert got_span_ids == expected_span_ids
@pytest.mark.parametrize("surface", ["filter", "select", "order"])
def test_traces_list_unknown_span_context_synthesizes(
signoz: types.SigNoz,