SceneSemanticsToolConfig
Field |
Type |
Repeated |
Description |
labels |
✅ |
Label set for zero-shot classification (each is embedded via the CLIP text encoder at instance creation; wrapped in "a photo of {label}") |
|
aggregation |
|||
sample_n |
|||
bucket_secs |
for BUCKETED; SDK defaults to 2.0 |
||
infer_every_ms |
min media-time between inferences; 0 = every frame |
||
model |
"" or "clip-vit-b32" (today’s model; named for forward compatibility) |