spark.read.format("geoparquet").load(None) in PySpark (which calls the JVM load() with no paths), .load() in Scala, and reading an existing directory that contains no files all fail with:
[UNABLE_TO_INFER_SCHEMA] Unable to infer schema for GeoParquet. It must be specified manually.
The message does not point at the actual cause, a missing or None path. That is easy to hit when an upstream step returns None for "no results" and the value is passed straight to load().
Spark removes the path option before it calls FileFormat.inferSchema and only passes the list of files it found, so the reader cannot tell a missing path from an empty directory: both arrive as an empty file list.
Proposal
When there are no input files, raise an AnalysisException whose message names both causes and the fixes: pass a path to GeoParquet files, or specify the schema. Reads with a user-specified schema, and globs that match nothing (which already fail with PATH_NOT_FOUND), stay as they are.
spark.read.format("geoparquet").load(None)in PySpark (which calls the JVMload()with no paths),.load()in Scala, and reading an existing directory that contains no files all fail with:The message does not point at the actual cause, a missing or
Nonepath. That is easy to hit when an upstream step returnsNonefor "no results" and the value is passed straight toload().Spark removes the
pathoption before it callsFileFormat.inferSchemaand only passes the list of files it found, so the reader cannot tell a missing path from an empty directory: both arrive as an empty file list.Proposal
When there are no input files, raise an
AnalysisExceptionwhose message names both causes and the fixes: pass a path to GeoParquet files, or specify the schema. Reads with a user-specified schema, and globs that match nothing (which already fail withPATH_NOT_FOUND), stay as they are.