Search & Discovery
The catalog is the front door to GeoLens. Every dataset you can see is reachable from search, and every search is reusable: filters drive a URL you can copy or save, the results page is the same regardless of role, and the bbox map is always live. This page walks through how to search effectively (keyword, faceted, semantic, and saved) so you can find what you need in a few keystrokes.
How catalog search works
Section titled “How catalog search works”GeoLens search is a single search box backed by three modes layered on the same catalog index:
- Keyword search: matches dataset title, description, keywords, contacts, lineage, and theme category, plus each record’s translated title and summary. Other metadata columns (license, organization, CRS, attribute values) aren’t part of the text match; filter on those in the rail instead. Default mode; results refresh as you type, about 300 ms after you stop.
- Faceted search: combines structured filters (keywords, record type, geometry type, location, organization, CRS, collection, dates) with the keyword query.
- Semantic search: when your administrator has enabled it, search runs against vector embeddings built from each dataset’s title, summary, keywords, and lineage, so a query like “rainfall over the Pacific Northwest” finds a dataset titled “WA precipitation monthly normals” even without keyword overlap.
The result set is the same shape in every mode: a paginated list of catalog records ordered by relevance. The URL fully captures your filter state, so you can paste a search link into chat or bookmark a query you run weekly.
Filter rail
Section titled “Filter rail”The left filter rail is where structured filters live. As you toggle filters, the result count next to each facet updates in place.
The most useful filters for narrowing a large catalog:
- Keywords: the keywords carried by each record. The rail lists the top 20 keywords in the current result set with a count beside each. Selecting more than one ANDs them: a dataset must carry every keyword you select.
- Type: Vector, Raster, Virtual Raster, or Table.
- Geometry type: restrict to
Point,LineString,Polygon,MultiPoint,MultiLineString, orMultiPolygon. It applies to vector datasets; use the Type control to narrow by record type instead. Useful when you know the layer type you need, e.g. finding only the road centerlines and not the road buffers. - Location: drag a rectangle on the inset map, or draw a polygon, and pick an Intersects or Within predicate. The shape is compared against each dataset’s spatial extent; only datasets whose extent matches appear. A rectangle is stored as four floats in the URL, so it’s reproducible.
- Organization: the dataset’s source organization.
- CRS: the dataset’s SRID.
- Collection: restrict to the members of one collection.
- Date Added and Temporal Extent: the record’s creation-date range, and
its OGC
datetimeinterval.
Combining filters
Section titled “Combining filters”Filters compose: each filter narrows the result set further. The active filter chips appear above the result list, and clicking the X on a chip removes that filter without clearing the rest. To clear everything at once, use Clear filters at the top of the rail.
Faceted search
Section titled “Faceted search”Faceted search is the combination of one or more filters with the keyword query. The result count badge by each facet shows how many datasets you’d see if you also enabled that facet: a “would-be” count, computed against the current search context, so you can preview the impact of a refinement.
For example, with the query roads and the Polygon geometry type already
selected, the Keywords facet might show:
transportation(12)infrastructure(8)urban(3)
Each number is the count of polygon datasets matching roads that also carry
that keyword. Clicking transportation re-runs the search with that filter
added.
URL persistence
Section titled “URL persistence”Every facet selection is captured in the URL: query string parameters like
?q=roads&geometry_type=POLYGON&keywords=transportation. Geometry types are
upper-case, and selecting several keywords repeats the parameter
(&keywords=transportation&keywords=urban). This means:
- You can bookmark a refined search and return to the same results on demand.
- Sharing a search link with a colleague delivers the exact filter context they need to see what you’re seeing.
- Filter changes rewrite the current URL in place rather than adding history entries, so Back leaves the catalog page instead of undoing your last refinement. Use Clear filters to start over.
Keyword and bbox filters live in the URL too. The bbox is stored as
?bbox=west,south,east,north (four decimal floats, EPSG:4326).
Semantic search
Section titled “Semantic search”Semantic search is an admin-toggled feature, enabled instance-wide by the
semantic_search_enabled setting (Admin -> Settings -> AI, not read from
.env). When it’s on and your query has text, GeoLens
blends keyword (full-text) results with vector-similarity results using
reciprocal-rank fusion. It doesn’t replace keyword matching with vector
matching; it merges the two rankings. If embeddings aren’t available (no
embeddings computed yet, or the provider errors), it falls back
gracefully to keyword-only results.
When semantic search is on, GeoLens compares your query against pre-computed vector embeddings built from every dataset’s title, summary, keywords, lineage, and any translated title or summary. The match is based on conceptual similarity, not literal token overlap, so:
- A query like
urban tree canopymay surface a dataset titled “MTL park-area green-cover survey 2023” even though no shared keywords appear. - A query like
water quality sampling siteswill favor datasets about monitoring stations even if those datasets don’t include the exact phrase.
Low-similarity matches are filtered out below an internal cutoff, so weak conceptual matches don’t crowd out stronger ones. Queries shorter than four characters skip the vector lookup altogether and return keyword results, so semantic ranking only kicks in once you’ve typed a real word. Results are ordered by relevance, blending embedding similarity with keyword matching.
When to use which mode
Section titled “When to use which mode”- Keyword is faster and more precise when you know the dataset name or a distinctive word from its description or keywords. Use this for everyday lookups.
- Faceted is best when you have a structural constraint (geometry type, keyword, bbox) but don’t know the exact dataset name.
- Semantic shines when you describe what you need in natural language and aren’t sure what naming conventions your team uses. Try it when keyword search returns nothing, or when you’re exploring an unfamiliar catalog.
Saved searches
Section titled “Saved searches”A saved search is a named, reusable search definition. Once saved, it appears
as a chip above the result list and re-runs against the live catalog every time
you open it, so a saved search for geometry_type=POLYGON plus
keywords=hydrology always returns the current matching set, not a frozen
snapshot.
To save a search:
- Run the search you want to preserve (filters, query, bbox, or whatever you’ve set up).
- Click Save Search in the filter toolbar. It appears once you have a query or a filter set, and only when you’re signed in.
- Give it a name (e.g., “Quarterly water-quality review”).
Saved searches are always per-user. They’re private to your account. There’s no visibility setting and no way to share a saved search with other users.
Saved searches differ from collections in two important ways:
- A saved search is a query: its membership changes as new datasets are
added that match the criteria. If you save
keywords=hydrologyand a new hydrology dataset is uploaded next month, it’s automatically in your saved search. - A collection is an explicit list: it contains exactly the datasets you added by hand, regardless of any catalog changes.
Use saved searches when the criteria matter; use collections when the specific datasets matter. See Collections for the explicit-membership alternative.
Recall and editing
Section titled “Recall and editing”Saved searches appear as chips above the result list on the catalog page, once you’re signed in. Click a chip to reload that query. To update a saved search, refine it in the search bar and save it again under the same name: that replaces the stored query. To remove one, click the X on its chip.
Programmatic equivalents
Section titled “Programmatic equivalents”The same catalog is available over OGC API Records: useful for clients that
need to drive search programmatically (a dashboard, a scheduled job, a custom
integration). Most UI filters have an API equivalent on the same route, though
not all of them go through CQL2: q, bbox, keywords, geometry_type,
srid, source_organization, record_type, collection_id, and datetime
are plain query parameters, while filter= (CQL2) covers title, description,
geometry_type, srid, source_organization, license, created, updated, and the
data-vintage dates. GET /api/collections/datasets/queryables is the
authoritative list.
For the full machine-client reference, see
OGC API & Standards Endpoints. It covers list, filter,
bbox, and CQL2 against /api/collections/datasets/items. The OGC API Records
endpoint is the canonical way to mirror the catalog into another tool, run a
nightly drift report, or build a custom search front-end.
The search bar’s query is q= on that same route. It runs the catalog text
search, and when semantic search is enabled it matches by meaning the way the
UI does. bbox= narrows the hits by each record’s extent rather than by its
features, and the response carries no relevance score; results arrive in
ranked order.
GeoLens also exposes its own search routes. /api/search/datasets/?q= runs the
identical query for the SDKs and the MCP server, and answers with the same OGC
Records FeatureCollection the items route returns. The difference is scope:
only the Records route reads type= and ids=, so pulling collection records
into a result set is a Records-route job. /api/search/facets/ has no Records
counterpart at all — it returns the counts that drive the filter rail, grouped
as record_type, keywords, source_organization, srid, and collections.
Since 1.14.1 both answer cross-origin browser requests: each sends a
credential-free Access-Control-Allow-Origin: * on an anonymous GET, so a
static page can call either route directly. A request that carries a credential
still needs its origin in CORS_ALLOWED_ORIGINS, as do the authenticated
/api/search/saved/ routes; see
Settings reference -> Network. The examples
repo’s
search/catalog.html
is that call from a browser, narrowed to the map view and drawn.
Saved searches over the API
Section titled “Saved searches over the API”Saved searches have a four-route CRUD surface under /api/search/saved/. Every
route needs an authenticated identity — a bearer token, or an X-Api-Key for a
key scoped full. A read_only key can list and read saved searches but is
refused on create and delete: that scope authenticates GET, HEAD, and
OPTIONS only, outside a short exemption list that does not include these
routes (see
API authentication -> API keys). Each identity
sees only its own saved searches:
GET /api/search/saved/lists them, most recently updated first, as{"searches": [...], "total": n}.skipandlimitpage it;limitdefaults to 50 and is capped at 200.POST /api/search/saved/takes{"name": ..., "params": {...}}and returns201.paramsis an opaque JSON object: the UI stores the catalog query parameters there and replays them when you click the chip.GET /api/search/saved/{search_id}returns one.DELETE /api/search/saved/{search_id}removes one,204on success.
There is no update route. POST upserts on the (user, name) pair, so posting a
name you already used replaces its stored params — that is the API behind
“save it again under the same name” above. A search_id belonging to someone
else reads as 404, not 403.
Tips and gotchas
Section titled “Tips and gotchas”- Empty results? Check the filter chips above the result list: a stale bbox or a forgotten keyword is the usual culprit. The no-results state offers Clear search & filters to start over.
- Wrong geometry? Some datasets have mixed geometry; the geometry-type filter matches the dataset’s primary geometry only. To find datasets with any of multiple types, leave the filter off and read the geometry type from each result card.
- Performance on large catalogs. Keyword search runs against GIN trigram
and full-text indexes, so it stays fast as the catalog grows; deep offset
pages are the expensive case. The bottleneck is usually the initial dataset
list render, not the search itself. Page size is a per-request
limit(default 10, capped at 200 on the search and Records routes), not an instance setting. - Semantic search and keywords. The stored embedding is built from the dataset’s title, summary, keywords, and lineage, so a keyword can influence a semantic match even when your query shares no words with the title. Structural constraints (geometry, CRS, extent, dates) are not part of the embedding; apply those in the filter rail.
See also
Section titled “See also”- Dataset detail: what each catalog hit looks like when you click into it
- Collections: group datasets by topic; the explicit-list alternative to saved searches
- OGC API & Standards Endpoints: programmatic search via OGC API Records
search/catalog.htmlin the examples repo: the same semantic search from a plain HTML page, with the map as the bbox filter- Settings reference: the admin-side semantic search toggle, embedding backfill, and other search configuration