SceneSemanticsToolConfig

Field

Type

Repeated

Description

labels

string

Label set for zero-shot classification (each is embedded via the CLIP text encoder at instance creation; wrapped in "a photo of {label}")

aggregation

SceneSemanticsAggregation

sample_n

int32

bucket_secs

double

for BUCKETED; SDK defaults to 2.0

infer_every_ms

double

min media-time between inferences; 0 = every frame

model

string

"" or "clip-vit-b32" (today’s model; named for forward compatibility)

Member of

Message

Description

LocalAnalysisConfig