Aller au contenu
getgeolens.com

Importing Data

Ce contenu n’est pas encore disponible dans votre langue.

Importing is how datasets get into GeoLens. The catalog supports five import modes (uploading a file, fetching one from a link, registering an existing PostGIS table, copying a layer from an external service, or referencing assets from a STAC API) and each mode handles its own validation, schema detection, and ingestion automatically. This page walks through each mode, the supported formats, how to create empty layers for collaborative editing, and how to re-import a dataset to update it without breaking its catalog identity.

The Import page is reached from Import Data in the Create menu in the top navigation bar. At the top of the page is a five-tab bar:

  • Upload File: push a file from your computer.
  • File URL: give GeoLens a link and let the server fetch the file.
  • Register Table: point at an existing PostGIS table on the same database.
  • Service URL: register a remote vector service URL (WFS, ArcGIS Feature Server, or OGC API Features).
  • STAC Catalog: ingest items from a STAC API endpoint (raster collections).

Each tab is a self-contained workflow. Every result is a dataset in the same catalog, but its storage and update behavior depends on the selected mode. See Where Your Data Lives for the managed-copy, in-place-table, and remote-reference differences.

Upload mode handles a file from your local filesystem. Drag a file (or files, for multi-file upload) onto the drop area, or click to browse.

Supported formats:

FormatExtensionNotes
Shapefile.zipRequired to be zipped (Shapefile is multi-file: .shp, .shx, .dbf, optional .prj)
File Geodatabase.zipA .zip containing a .gdb directory. GeoLens tells the two zip container formats apart automatically; a FileGDB can fan out to one dataset per layer
GeoPackage.gpkgVector or tabular layers; a multi-layer GeoPackage can fan out to one dataset per layer. Raster GeoPackages aren’t supported — upload rasters as GeoTIFF
GeoJSON.geojson, .jsonSingle FeatureCollection per file
FlatGeobuf.fgbSingle-file vector with a built-in spatial index; no zipping needed
KML / KMZ.kml, .kmzGoogle Earth vector. A KMZ is a zipped KML — one format in two containers — so a KMZ upload is recorded as KML
GeoParquet.parquetGeoParquet 1.0/1.1 with WKB geometry; a plain Parquet file without geo metadata imports as a non-spatial table. The bridge from any warehouse that writes Parquet (BigQuery, Snowflake, Redshift, Databricks, DuckDB)
CSV with geometry.csvGeometry as WKT in a column, or lat/lon columns; auto-detected
Excel.xlsx, .xlsTabular; geometry as WKT in a column, or lat/lon columns
GeoTIFF.tif, .tiffRaster

The maximum upload size is governed by your instance’s UPLOAD_MAX_SIZE_MB setting (default 500 MB). For files larger than the limit, ask your admin to raise the limit, or bring the data in another way: Register mode for a PostGIS table, Service mode for a service URL. See Configuration reference for the admin-side knob. Two other limits apply: a batch is capped at 25 files (fewer when you’re close to your dataset quota), and an upload that would push you past that quota is rejected with a Dataset limit reached banner.

A reverse proxy or hosting provider can impose a lower request-size limit than GeoLens. If it rejects a file that your instance permits, ask your admin about that limit or use File URL mode when the file is available at a direct download URL.

An upload runs in stages. GeoLens uploads the file, detects its geometry, schema, and CRS, and shows a preview. You then review the metadata — name, description, visibility, a CRS override, the geometry column for CSV and spreadsheet sources, and compression, resampling, and nodata for rasters — and commit. For a multi-layer container (a GeoPackage or a File Geodatabase), the review step also shows a Layer picker — layers, not sheets; the sheet vocabulary is reserved for spreadsheets — plus an Ingest all layers as separate datasets option that lands every layer as its own dataset. The ingest itself runs as a tracked job.

The job reports one of six statuses: Pending, Running, Complete, Failed, Cancelled, or Fanned out, alongside a per-step readout (validating, loading features, finalizing, and for rasters COG conversion and preview generation). Fanned out is the multi-layer-source case (most often a GeoPackage or File Geodatabase, but any OGR source with more than one layer — including a .zip holding several Shapefiles): each layer lands as its own dataset and a results panel lists which layers succeeded and which failed. Cancelled means nothing was ever attempted — an upload that was requested but never completed settles there instead of looking like a broken ingest. A failed job offers a Retry.

An import that completes with caveats shows an Import completed with warnings banner, first on the upload success screen and then on the dataset page. See Import warnings for what each one means.

File URL mode takes a direct HTTP or HTTPS link to one file and has the server download it, so the bytes never pass through your browser. The supported formats and the size limit are the ones Upload mode lists, and the download also stops early where your remaining storage quota is the smaller ceiling. The link has to resolve to the file itself rather than to a landing page. Where the URL path does not end in the file name, as with a download link keyed by an id, fill in the optional Filename override so GeoLens can tell what it is receiving.

The download runs as a background job, so a large file or a slow origin is not bounded by your instance’s proxy read timeout. While it runs you get the progress bar, step label and Cancel and start over control that any other job has. Your instance caps a single download with URL_IMPORT_FETCH_MAX_SECONDS, which defaults to 30 minutes; see Configuration reference.

You can switch to another tab while the download runs and come back to it. Reloading the page cannot: the job carries on server-side, but the page has no way to re-attach to it. Once the file has landed, the preview, metadata review and commit steps are the ones Upload mode describes.

Register mode tells GeoLens about an existing PostGIS table on the same database that the API connects to. No data copy happens: the catalog records a reference to the table, and queries against the dataset go directly against the underlying table.

Use Register mode when:

  • A nightly ETL pipeline produces a PostGIS table outside of GeoLens, and you want the table catalogued without copying the rows.
  • A team is editing a PostGIS table directly through QGIS or another client, and the table should appear in the catalog without round-trips.
  • The dataset is too large to upload but is already in PostGIS.

The form doesn’t ask for a schema or table name. GeoLens discovers tables that aren’t registered yet and lists them as a searchable pick-list, each row showing its geometry type and an estimated row count. Pick one, choose Private or Public, and register it. Once registered, the dataset behaves like any other catalog dataset: it’s searchable, exportable, mappable. Edits to the underlying table appear in the catalog the next time the dataset is queried.

Discovery only looks inside the GeoLens data schema (data, or data_t_{tid} on a multi-tenant instance), and only at base tables — a table your ETL wrote to public or another schema won’t appear in the list, and neither will a view. It also skips names ending in _staging or _old, skips spatial_ref_sys, and returns at most 1,000 tables, so a very large instance may not list everything at once.

Service mode registers an external service URL. GeoLens probes the URL to detect the service type automatically. The supported types are:

  • WFS: OGC Web Feature Service (standard URLs containing /wfs).
  • OGC API Features: newer OGC API standard (URLs containing /collections or matching the OGC API Common landing page shape).
  • ArcGIS Feature Server: Esri Feature Server REST endpoint (URLs containing /FeatureServer or /MapServer).

After auto-detection, the form lists the service’s available layers; pick one to import as a GeoLens dataset. GeoLens fetches the layer into a local PostGIS table and stores a system-managed pointer to the service and layer for manual refresh. Feature queries and tiles use the local copy rather than proxying through to the service at request time.

Authenticating a protected service. The credential form depends on what the URL looks like.

For WFS and OGC API Features, the Authentication select offers four methods: No authentication, Bearer token, Username and password, or API key in a header. A bearer token is sent as an Authorization header; username and password go as HTTP Basic auth; a header credential goes under whatever header name the provider specifies — X-API-Key, Ocp-Apim-Subscription-Key, and key are common examples.

For an ArcGIS Feature Server or Map Server, the same select offers No authentication, Sign in with username and password, or Paste a token or API key:

  • Sign in exchanges your ArcGIS Online or Portal for ArcGIS username and password for a token, the same way signing into the portal in a browser would. The Portal URL field pre-fills for an ArcGIS Online service URL; for Portal for ArcGIS, enter your organization’s portal URL yourself. GeoLens holds the password only for that one sign-in request and clears it the instant the attempt settles, success or failure. The minted token is valid for 60 minutes, so an import that runs longer needs a fresh sign-in. Repeated sign-in attempts are rate-limited, both per GeoLens account and per ArcGIS account — retype a mistyped password rather than retrying it, since ArcGIS itself locks an account out after five failed sign-ins in fifteen minutes. Sign-in reaches the portal directly, so it can’t be used against a portal on a private network; paste a token instead in that case.
  • Paste a token or API key is the original single-field flow: generate an API key from your portal’s Sharing API with client=referer if your account can create one (Viewer accounts and accounts using single sign-on or multi-factor authentication usually can’t). A token lasts at most 15 days.

Either way, the credential is scoped to that one import: it travels with the job for that import only and is purged once the job reaches a terminal state, the same as it always was for an ArcGIS token. It is not part of the dataset’s stored origin pointer, so a later Refresh from source (see the Sources tab) of a protected source asks for a credential again — the dialog offers the same authentication choices described above. Use a short-lived, least-privilege credential, and never embed one in the source URL itself.

A few notes:

  • The service URL should be the service root, not a feature-list URL. GeoLens probes for the capabilities document.
  • For OGC API Features specifically, the root is the URL ending in /api/ (or wherever the service exposes its landing document).
  • ArcGIS Feature Servers can be either ArcGIS Online or ArcGIS Enterprise; both work with the same auto-detection.

STAC mode ingests STAC items from a STAC API endpoint. Use it to pull raster collections (imagery, elevation, classification grids) from external STAC catalogs into GeoLens.

The flow is:

  1. Paste the STAC API root (e.g., https://example.com/stac/) and click Connect.
  2. Pick a collection from the discovered list.
  3. GeoLens lists that collection’s first 50 items, each with its date, EPSG code, GSD, cloud cover, and asset count. Items with no COG asset are greyed out and can’t be selected.
  4. Tick the items you want. Scoping is per-item — there are no bbox, datetime, or limit filters.
  5. A Review download size step totals the selected assets’ sizes, taken from the STAC manifest, before you confirm.

GeoLens registers each matching STAC item as a raster dataset in the catalog; the items become visible in search and addable to maps. The asset URLs are stored as references. No asset data is copied unless you also export the dataset.

From 1.19.0 a catalog that requires a credential can be imported and refreshed. The Authentication select on the STAC tab offers the same four methods the WFS and OGC API Features form does: No authentication, Bearer token, Username and password, or API key in a header under whatever header name the provider specifies. Every read GeoLens makes against the catalog while you browse carries the credential. The refresh dialog on the dataset page offers the same methods, so a catalog you imported with one can be refreshed with one.

Two consequences are worth knowing before you start:

  • The credential reaches the catalog’s own origin and nothing else. A STAC document can point at another host, and GeoLens reads such an address anonymously rather than forwarding your key to it.
  • An asset that needs its own credential cannot be tiled. Tiles are rendered by fetching the asset URL out of process, with no credential attached, so a protected asset renders nothing and a refresh reports it as inaccessible. A protected catalog whose assets are public works normally.

A refresh of a catalog that needed a credential on its last successful run is refused up front when you supply none, rather than dispatched to collect a 401. Credentials are request-only: GeoLens never stores one between runs, so every refresh asks for it again. Passing one to a refresh also needs the instance’s shared credential store, the same requirement service refreshes have; see Where Your Data Lives.

Some imports succeed but change your data on the way in. When that happens the ingest job records structured warnings and the Import completed with warnings banner appears — on the upload success screen while you’re still watching the job, and then permanently on the dataset detail page. The dataset page reads the dataset’s most recent ingest job and only shows the banner once that job has completed, so a re-import’s warnings replace the original import’s. Datasets registered from an existing PostGIS table have no ingest job, so they never carry one.

The banner can hold five kinds of warning. The name in parentheses is what to quote in a support thread: for the first three it’s the kind value on an entry in the job record’s warnings array, and for the last two it’s a field on the job record itself.

  • Reserved column names renamed (reserved_rename). A source column arrived named gid, geom, geometry, geom_4326, fid, or ogc_fid, or with a : in its name (Socrata exports ship :id, :created_at). GeoLens prefixes reserved names with src_ and converts colon names to letters, digits, and underscores. This keeps the pipeline’s geometry and key columns available. Update styles, filters, and API queries to use the new names. The banner lists every original -> renamed pair.
  • Shapefile field name collisions (dbf_truncation_collision). Checked on .zip uploads only. Two or more source field names share their first 10 characters (compared case-insensitively) — the DBF field-name cap — so they may have been merged into one column during import. Check the imported schema, and if a field is missing, rename the source fields so their first 10 characters differ and re-upload.
  • Geometry outside the Web Mercator bounds (mercator_clip). Vector geometry ran past longitude ±180° or latitude ±85.06°, the envelope GeoLens’s tiles cover. The rows are still there; only their geometry changed, and the banner counts features trimmed at the boundary separately from features that lost their geometry entirely. Usually the source has swapped coordinate order (lat/lon instead of lon/lat) or the wrong CRS — check both before re-importing. This warning has a second form that says the check was skipped: when the source CRS has a narrow area of validity the bounds can’t be expressed in it, so nothing was removed, but out-of-range geometry may still be present and won’t render on a map.
  • Original file archive failed (archive_failed). The import itself worked, but copying your original upload to durable storage didn’t. The dataset is complete and queryable; what’s missing is the archived source file. Neither the banner nor the API tells you why the copy failed; the reason is recorded in the server logs and on the job record, so an admin has to look. Re-upload if you need the original kept, and tell your admin.
  • Temporal fields dropped (temporal_parse_errors). Raster imports only. A start or end date you entered on the review step wasn’t valid ISO-8601, so it was dropped rather than guessed at. The banner shows the raw value; re-enter it on the dataset’s metadata editor.

For scripted checks, the same payload is on GET /api/jobs/by-dataset/{dataset_id}, which returns the dataset’s latest ingest job (or null when it has none) with a warnings array plus the archive_failed and temporal_parse_errors fields.

For collaborative data entry, where a team will edit features over time through a GIS client, you can create an empty layer instead of uploading a file. This isn’t part of the Import page: use the Create menu in the top navigation bar and choose Dataset. The dialog prompts for a title and a schema (column names + types), then creates an empty PostGIS table (in EPSG:4326) backed by a new catalog entry.

Use empty layers when:

  • A team is digitizing features over time (e.g., a survey crew adding points as they’re collected).
  • The dataset’s schema is known but the data hasn’t been collected yet.
  • A workflow uses GeoLens as the database of record but creates rows programmatically through the API or directly in PostGIS.

Once the empty layer exists, it’s the same as any other dataset, except the data tab shows zero rows until edits arrive.

Re-import is how you replace a dataset’s data without breaking its catalog identity. From the existing dataset’s detail page, click Re-Upload in the page header. That opens a dialog on the dataset page — you never leave it — which runs its own flow: pick a file or a service, preview the result — for a vector dataset that preview includes a schema diff against the current dataset — commit, and track the job. The file picker accepts the same formats as Upload mode, including the 1.16 additions (FlatGeobuf, KML/KMZ, and a zipped File Geodatabase). Raster re-uploads take a replacement file only: there’s no service option and no schema comparison, and confirming starts a COG conversion that can run for several minutes. VRT datasets don’t have this action at all; they use the regenerate flow instead.

What re-import preserves:

  • Dataset ID and URL: bookmarks and shared links keep working.
  • Metadata: title, description, tags, custom fields are unchanged.
  • Permissions: visibility and access list are preserved.
  • Comments and audit trail: the dataset’s history continues, with the re-import logged as a new entry on the sources tab.

What re-import replaces:

  • Feature rows: the entire dataset is replaced with the new file’s contents.
  • Schema: if the new file has different columns, the schema updates; this can break downstream maps that reference removed columns, so re-importing with a changed schema warrants a heads-up to consumers.
  • Spatial extent: recomputed from the new feature geometries.

What re-import refuses, from 1.19.1:

  • A replacement with no geometry over a dataset that stores geometry. A file with no coordinates, such as a plain CSV, would leave the dataset a plain table, so the preview stops with a message saying nothing was changed and that the file belongs in a new dataset. An API client that skips the preview and commits anyway is accepted, and its job then fails with the same message, leaving the dataset as it was. A service replacement whose layer arrives without geometry fails its job the same way.
  • A Parquet file over the ingest ceiling on rows or cells. The preview reports the ceiling and the file’s count, as the first-time import preview does, instead of calling the file malformed.

The same re-import action works across modes: a dataset originally created from a file upload can be re-imported from a service URL, or vice versa. The dataset’s identity is independent of the import method.

A common question: when should a dataset be registered (service or PostGIS) vs. uploaded?

ChooseWhen
UploadOne-off datasets; data won’t change; small to medium size; you want GeoLens to fully own the storage
Register PostGISData is already in PostGIS; updates happen outside GeoLens; large size; PostGIS should stay authoritative
Import serviceData lives in an external feature service; you want a local GeoLens copy for querying and tiling, with manual re-fetch from the same source
STAC ingestRaster collections from a STAC catalog; you want pointers, not copies
Create emptyForward-looking: a team will populate the layer over time through GeoLens or external clients

All five import modes require the editor role. The role is checked both client-side (the Create menu omits Import Data) and server-side (the API rejects import requests from non-editors). For a refresher on GeoLens’s role model, see User management & RBAC.

The owner of a newly imported dataset is the user who imported it, regardless of mode. New datasets are private unless you say otherwise: the upload and register forms offer a Private/Public picker that starts on Private, and STAC imports land private. You can change visibility later to private, internal, or public from the access tab on the dataset detail page.

The same import flows work over the API for headless automation. Use the same dataset endpoints and pass file content as a multipart body, or register a service URL via JSON. For authentication, you’ll need an API key or a JWT. See API Authentication for the auth options.

  • Shapefile encoding. Shapefiles use a .cpg file (or system default) for character encoding. If your .dbf has non-ASCII text and no .cpg, GeoLens defaults to UTF-8; uploads with mojibake usually mean the source was Latin-1 or Windows-1252. Convert the encoding before zipping, or include an explicit .cpg file.
  • CSV geometry. GeoLens auto-detects WKT in a column named wkt, geom, geometry, the_geom, or shape, and auto-detects latitude (lat, latitude, y, lat_dd, ycoord) / longitude (lon, lng, long, longitude, x, lon_dd, xcoord) pairs. Auto-detection is only the default: for CSV and spreadsheet uploads the review step shows a geometry section where you pick the lat/lng pair or the WKT column yourself from the file’s own columns, so differently named columns need no renaming.
  • CRS declaration. Always include a .prj file in zipped Shapefiles and a CRS declaration in GeoJSON files when possible. GeoLens defaults to EPSG:4326 if no CRS is declared, which is wrong for projected data and breaks downstream styling. If a file arrives without one, set the correct code in the review step’s CRS Override field before you commit. From 1.19.1 the field takes only an EPSG code the instance knows: a code that neither PROJ nor PostGIS has is refused when you commit, with a message naming it. Leave the field empty to keep the source’s CRS.
  • Re-import schema changes. A re-import replaces the schema: a column missing from the new file is dropped from the dataset. The re-import dialog shows a schema diff before you commit, so check it. To keep an old column the new file doesn’t have, re-import with that column included (even if empty).