Skip to content
getgeolens.com

Search & Discovery

The catalog is the front door to GeoLens. Every dataset you can see is reachable from search, and every search is reusable: filters drive a URL you can copy or save, the results page is the same regardless of role, and the bbox map is always live. This page walks through how to search effectively (keyword, faceted, semantic, and saved) so you can find what you need in a few keystrokes.

GeoLens search is a single search box backed by three modes layered on the same catalog index:

  • Keyword search: matches dataset title, description, keywords, contacts, lineage, and theme category, plus each record’s translated title and summary. Other metadata columns (license, organization, CRS, attribute values) aren’t part of the text match; filter on those in the rail instead. Default mode; results refresh as you type, about 300 ms after you stop.
  • Faceted search: combines structured filters (keywords, record type, geometry type, location, organization, CRS, collection, dates) with the keyword query.
  • Semantic search: when your administrator has enabled it, search runs against vector embeddings built from each dataset’s title, summary, keywords, and lineage, so a query like “rainfall over the Pacific Northwest” finds a dataset titled “WA precipitation monthly normals” even without keyword overlap.

The result set is the same shape in every mode: a paginated list of catalog records ordered by relevance. The URL fully captures your filter state, so you can paste a search link into chat or bookmark a query you run weekly.

The left filter rail is where structured filters live. As you toggle filters, the result count next to each facet updates in place.

The most useful filters for narrowing a large catalog:

  • Keywords: the keywords carried by each record. The rail lists the top 20 keywords in the current result set with a count beside each. Selecting more than one ANDs them: a dataset must carry every keyword you select.
  • Type: Vector, Raster, Virtual Raster, or Table.
  • Geometry type: restrict to Point, LineString, Polygon, MultiPoint, MultiLineString, or MultiPolygon. It applies to vector datasets; use the Type control to narrow by record type instead. Useful when you know the layer type you need, e.g. finding only the road centerlines and not the road buffers.
  • Location: drag a rectangle on the inset map, or draw a polygon, and pick an Intersects or Within predicate. The shape is compared against each dataset’s spatial extent; only datasets whose extent matches appear. A rectangle is stored as four floats in the URL, so it’s reproducible.
  • Organization: the dataset’s source organization.
  • CRS: the dataset’s SRID.
  • Collection: restrict to the members of one collection.
  • Date Added and Temporal Extent: the record’s creation-date range, and its OGC datetime interval.

Filters compose: each filter narrows the result set further. The active filter chips appear above the result list, and clicking the X on a chip removes that filter without clearing the rest. To clear everything at once, use Clear filters at the top of the rail.

Faceted search is the combination of one or more filters with the keyword query. The result count badge by each facet shows how many datasets you’d see if you also enabled that facet: a “would-be” count, computed against the current search context, so you can preview the impact of a refinement.

For example, with the query roads and the Polygon geometry type already selected, the Keywords facet might show:

  • transportation (12)
  • infrastructure (8)
  • urban (3)

Each number is the count of polygon datasets matching roads that also carry that keyword. Clicking transportation re-runs the search with that filter added.

Every facet selection is captured in the URL: query string parameters like ?q=roads&geometry_type=POLYGON&keywords=transportation. Geometry types are upper-case, and selecting several keywords repeats the parameter (&keywords=transportation&keywords=urban). This means:

  • You can bookmark a refined search and return to the same results on demand.
  • Sharing a search link with a colleague delivers the exact filter context they need to see what you’re seeing.
  • Filter changes rewrite the current URL in place rather than adding history entries, so Back leaves the catalog page instead of undoing your last refinement. Use Clear filters to start over.

Keyword and bbox filters live in the URL too. The bbox is stored as ?bbox=west,south,east,north (four decimal floats, EPSG:4326).

Semantic search is an admin-toggled feature, enabled instance-wide by the semantic_search_enabled setting (Admin -> Settings -> AI, not read from .env). When it’s on and your query has text, GeoLens blends keyword (full-text) results with vector-similarity results using reciprocal-rank fusion. It doesn’t replace keyword matching with vector matching; it merges the two rankings. If embeddings aren’t available (no embeddings computed yet, or the provider errors), it falls back gracefully to keyword-only results.

When semantic search is on, GeoLens compares your query against pre-computed vector embeddings built from every dataset’s title, summary, keywords, lineage, and any translated title or summary. The match is based on conceptual similarity, not literal token overlap, so:

  • A query like urban tree canopy may surface a dataset titled “MTL park-area green-cover survey 2023” even though no shared keywords appear.
  • A query like water quality sampling sites will favor datasets about monitoring stations even if those datasets don’t include the exact phrase.

Low-similarity matches are filtered out below an internal cutoff, so weak conceptual matches don’t crowd out stronger ones. Queries shorter than four characters skip the vector lookup altogether and return keyword results, so semantic ranking only kicks in once you’ve typed a real word. Results are ordered by relevance, blending embedding similarity with keyword matching.

  • Keyword is faster and more precise when you know the dataset name or a distinctive word from its description or keywords. Use this for everyday lookups.
  • Faceted is best when you have a structural constraint (geometry type, keyword, bbox) but don’t know the exact dataset name.
  • Semantic shines when you describe what you need in natural language and aren’t sure what naming conventions your team uses. Try it when keyword search returns nothing, or when you’re exploring an unfamiliar catalog.

A saved search is a named, reusable search definition. Once saved, it appears as a chip above the result list and re-runs against the live catalog every time you open it, so a saved search for geometry_type=POLYGON plus keywords=hydrology always returns the current matching set, not a frozen snapshot.

To save a search:

  1. Run the search you want to preserve (filters, query, bbox, or whatever you’ve set up).
  2. Click Save Search in the filter toolbar. It appears once you have a query or a filter set, and only when you’re signed in.
  3. Give it a name (e.g., “Quarterly water-quality review”).

Saved searches are always per-user. They’re private to your account. There’s no visibility setting and no way to share a saved search with other users.

Saved searches differ from collections in two important ways:

  • A saved search is a query: its membership changes as new datasets are added that match the criteria. If you save keywords=hydrology and a new hydrology dataset is uploaded next month, it’s automatically in your saved search.
  • A collection is an explicit list: it contains exactly the datasets you added by hand, regardless of any catalog changes.

Use saved searches when the criteria matter; use collections when the specific datasets matter. See Collections for the explicit-membership alternative.

Saved searches appear as chips above the result list on the catalog page, once you’re signed in. Click a chip to reload that query. To update a saved search, refine it in the search bar and save it again under the same name: that replaces the stored query. To remove one, click the X on its chip.

The same catalog is available over OGC API Records: useful for clients that need to drive search programmatically (a dashboard, a scheduled job, a custom integration). Most UI filters have an API equivalent on the same route, though not all of them go through CQL2: q, bbox, keywords, geometry_type, srid, source_organization, record_type, collection_id, and datetime are plain query parameters, while filter= (CQL2) covers title, description, geometry_type, srid, source_organization, license, created, updated, and the data-vintage dates. GET /api/collections/datasets/queryables is the authoritative list.

For the full machine-client reference, see OGC API & Standards Endpoints. It covers list, filter, bbox, and CQL2 against /api/collections/datasets/items. The OGC API Records endpoint is the canonical way to mirror the catalog into another tool, run a nightly drift report, or build a custom search front-end.

The search bar’s query is q= on that same route. It runs the catalog text search, and when semantic search is enabled it matches by meaning the way the UI does. bbox= narrows the hits by each record’s extent rather than by its features, and the response carries no relevance score; results arrive in ranked order.

GeoLens also exposes its own search routes. /api/search/datasets/?q= runs the identical query for the SDKs and the MCP server, and answers with the same OGC Records FeatureCollection the items route returns. The difference is scope: only the Records route reads type= and ids=, so pulling collection records into a result set is a Records-route job. /api/search/facets/ has no Records counterpart at all — it returns the counts that drive the filter rail, grouped as record_type, keywords, source_organization, srid, and collections.

Since 1.14.1 both answer cross-origin browser requests: each sends a credential-free Access-Control-Allow-Origin: * on an anonymous GET, so a static page can call either route directly. A request that carries a credential still needs its origin in CORS_ALLOWED_ORIGINS, as do the authenticated /api/search/saved/ routes; see Settings reference -> Network. The examples repo’s search/catalog.html is that call from a browser, narrowed to the map view and drawn.

Saved searches have a four-route CRUD surface under /api/search/saved/. Every route needs an authenticated identity — a bearer token, or an X-Api-Key for a key scoped full. A read_only key can list and read saved searches but is refused on create and delete: that scope authenticates GET, HEAD, and OPTIONS only, outside a short exemption list that does not include these routes (see API authentication -> API keys). Each identity sees only its own saved searches:

  • GET /api/search/saved/ lists them, most recently updated first, as {"searches": [...], "total": n}. skip and limit page it; limit defaults to 50 and is capped at 200.
  • POST /api/search/saved/ takes {"name": ..., "params": {...}} and returns 201. params is an opaque JSON object: the UI stores the catalog query parameters there and replays them when you click the chip.
  • GET /api/search/saved/{search_id} returns one.
  • DELETE /api/search/saved/{search_id} removes one, 204 on success.

There is no update route. POST upserts on the (user, name) pair, so posting a name you already used replaces its stored params — that is the API behind “save it again under the same name” above. A search_id belonging to someone else reads as 404, not 403.

  • Empty results? Check the filter chips above the result list: a stale bbox or a forgotten keyword is the usual culprit. The no-results state offers Clear search & filters to start over.
  • Wrong geometry? Some datasets have mixed geometry; the geometry-type filter matches the dataset’s primary geometry only. To find datasets with any of multiple types, leave the filter off and read the geometry type from each result card.
  • Performance on large catalogs. Keyword search runs against GIN trigram and full-text indexes, so it stays fast as the catalog grows; deep offset pages are the expensive case. The bottleneck is usually the initial dataset list render, not the search itself. Page size is a per-request limit (default 10, capped at 200 on the search and Records routes), not an instance setting.
  • Semantic search and keywords. The stored embedding is built from the dataset’s title, summary, keywords, and lineage, so a keyword can influence a semantic match even when your query shares no words with the title. Structural constraints (geometry, CRS, extent, dates) are not part of the embedding; apply those in the filter rail.
  • Dataset detail: what each catalog hit looks like when you click into it
  • Collections: group datasets by topic; the explicit-list alternative to saved searches
  • OGC API & Standards Endpoints: programmatic search via OGC API Records
  • search/catalog.html in the examples repo: the same semantic search from a plain HTML page, with the map as the bbox filter
  • Settings reference: the admin-side semantic search toggle, embedding backfill, and other search configuration