JSON Files

Spark SQL can automatically infer the schema of a JSON dataset and load it as a Dataset[Row]. This conversion can be done using SparkSession.read.json() on either a Dataset[String], or a JSON file.

Note that the file that is offered as a json file is not a typical JSON file. Each line must contain a separate, self-contained valid JSON object. For more information, please see JSON Lines text format, also called newline-delimited JSON.

For a regular multi-line JSON file, set the multiLine option to true.

  1. // Primitive types (Int, String, etc) and Product types (case classes) encoders are
  2. // supported by importing this when creating a Dataset.
  3. import spark.implicits._
  4. // A JSON dataset is pointed to by path.
  5. // The path can be either a single text file or a directory storing text files
  6. val path = "examples/src/main/resources/people.json"
  7. val peopleDF = spark.read.json(path)
  8. // The inferred schema can be visualized using the printSchema() method
  9. peopleDF.printSchema()
  10. // root
  11. // |-- age: long (nullable = true)
  12. // |-- name: string (nullable = true)
  13. // Creates a temporary view using the DataFrame
  14. peopleDF.createOrReplaceTempView("people")
  15. // SQL statements can be run by using the sql methods provided by spark
  16. val teenagerNamesDF = spark.sql("SELECT name FROM people WHERE age BETWEEN 13 AND 19")
  17. teenagerNamesDF.show()
  18. // +------+
  19. // | name|
  20. // +------+
  21. // |Justin|
  22. // +------+
  23. // Alternatively, a DataFrame can be created for a JSON dataset represented by
  24. // a Dataset[String] storing one JSON object per string
  25. val otherPeopleDataset = spark.createDataset(
  26. """{"name":"Yin","address":{"city":"Columbus","state":"Ohio"}}""" :: Nil)
  27. val otherPeople = spark.read.json(otherPeopleDataset)
  28. otherPeople.show()
  29. // +---------------+----+
  30. // | address|name|
  31. // +---------------+----+
  32. // |[Columbus,Ohio]| Yin|
  33. // +---------------+----+

Find full example code at “examples/src/main/scala/org/apache/spark/examples/sql/SQLDataSourceExample.scala” in the Spark repo.

Spark SQL can automatically infer the schema of a JSON dataset and load it as a Dataset<Row>. This conversion can be done using SparkSession.read().json() on either a Dataset<String>, or a JSON file.

Note that the file that is offered as a json file is not a typical JSON file. Each line must contain a separate, self-contained valid JSON object. For more information, please see JSON Lines text format, also called newline-delimited JSON.

For a regular multi-line JSON file, set the multiLine option to true.

  1. import org.apache.spark.sql.Dataset;
  2. import org.apache.spark.sql.Row;
  3. // A JSON dataset is pointed to by path.
  4. // The path can be either a single text file or a directory storing text files
  5. Dataset<Row> people = spark.read().json("examples/src/main/resources/people.json");
  6. // The inferred schema can be visualized using the printSchema() method
  7. people.printSchema();
  8. // root
  9. // |-- age: long (nullable = true)
  10. // |-- name: string (nullable = true)
  11. // Creates a temporary view using the DataFrame
  12. people.createOrReplaceTempView("people");
  13. // SQL statements can be run by using the sql methods provided by spark
  14. Dataset<Row> namesDF = spark.sql("SELECT name FROM people WHERE age BETWEEN 13 AND 19");
  15. namesDF.show();
  16. // +------+
  17. // | name|
  18. // +------+
  19. // |Justin|
  20. // +------+
  21. // Alternatively, a DataFrame can be created for a JSON dataset represented by
  22. // a Dataset<String> storing one JSON object per string.
  23. List<String> jsonData = Arrays.asList(
  24. "{\"name\":\"Yin\",\"address\":{\"city\":\"Columbus\",\"state\":\"Ohio\"}}");
  25. Dataset<String> anotherPeopleDataset = spark.createDataset(jsonData, Encoders.STRING());
  26. Dataset<Row> anotherPeople = spark.read().json(anotherPeopleDataset);
  27. anotherPeople.show();
  28. // +---------------+----+
  29. // | address|name|
  30. // +---------------+----+
  31. // |[Columbus,Ohio]| Yin|
  32. // +---------------+----+

Find full example code at “examples/src/main/java/org/apache/spark/examples/sql/JavaSQLDataSourceExample.java” in the Spark repo.

Spark SQL can automatically infer the schema of a JSON dataset and load it as a DataFrame. This conversion can be done using SparkSession.read.json on a JSON file.

Note that the file that is offered as a json file is not a typical JSON file. Each line must contain a separate, self-contained valid JSON object. For more information, please see JSON Lines text format, also called newline-delimited JSON.

For a regular multi-line JSON file, set the multiLine parameter to True.

  1. # spark is from the previous example.
  2. sc = spark.sparkContext
  3. # A JSON dataset is pointed to by path.
  4. # The path can be either a single text file or a directory storing text files
  5. path = "examples/src/main/resources/people.json"
  6. peopleDF = spark.read.json(path)
  7. # The inferred schema can be visualized using the printSchema() method
  8. peopleDF.printSchema()
  9. # root
  10. # |-- age: long (nullable = true)
  11. # |-- name: string (nullable = true)
  12. # Creates a temporary view using the DataFrame
  13. peopleDF.createOrReplaceTempView("people")
  14. # SQL statements can be run by using the sql methods provided by spark
  15. teenagerNamesDF = spark.sql("SELECT name FROM people WHERE age BETWEEN 13 AND 19")
  16. teenagerNamesDF.show()
  17. # +------+
  18. # | name|
  19. # +------+
  20. # |Justin|
  21. # +------+
  22. # Alternatively, a DataFrame can be created for a JSON dataset represented by
  23. # an RDD[String] storing one JSON object per string
  24. jsonStrings = ['{"name":"Yin","address":{"city":"Columbus","state":"Ohio"}}']
  25. otherPeopleRDD = sc.parallelize(jsonStrings)
  26. otherPeople = spark.read.json(otherPeopleRDD)
  27. otherPeople.show()
  28. # +---------------+----+
  29. # | address|name|
  30. # +---------------+----+
  31. # |[Columbus,Ohio]| Yin|
  32. # +---------------+----+

Find full example code at “examples/src/main/python/sql/datasource.py” in the Spark repo.

Spark SQL can automatically infer the schema of a JSON dataset and load it as a DataFrame. using the read.json() function, which loads data from a directory of JSON files where each line of the files is a JSON object.

Note that the file that is offered as a json file is not a typical JSON file. Each line must contain a separate, self-contained valid JSON object. For more information, please see JSON Lines text format, also called newline-delimited JSON.

For a regular multi-line JSON file, set a named parameter multiLine to TRUE.

  1. # A JSON dataset is pointed to by path.
  2. # The path can be either a single text file or a directory storing text files.
  3. path <- "examples/src/main/resources/people.json"
  4. # Create a DataFrame from the file(s) pointed to by path
  5. people <- read.json(path)
  6. # The inferred schema can be visualized using the printSchema() method.
  7. printSchema(people)
  8. ## root
  9. ## |-- age: long (nullable = true)
  10. ## |-- name: string (nullable = true)
  11. # Register this DataFrame as a table.
  12. createOrReplaceTempView(people, "people")
  13. # SQL statements can be run by using the sql methods.
  14. teenagers <- sql("SELECT name FROM people WHERE age >= 13 AND age <= 19")
  15. head(teenagers)
  16. ## name
  17. ## 1 Justin

Find full example code at “examples/src/main/r/RSparkSQLExample.R” in the Spark repo.

  1. CREATE TEMPORARY VIEW jsonTable
  2. USING org.apache.spark.sql.json
  3. OPTIONS (
  4. path "examples/src/main/resources/people.json"
  5. )
  6. SELECT * FROM jsonTable

Data Source Option

Data source options of JSON can be set via:

  • the .option/.options methods of
    • DataFrameReader
    • DataFrameWriter
    • DataStreamReader
    • DataStreamWriter
  • the built-in functions below
    • from_json
    • to_json
    • schema_of_json
  • OPTIONS clause at CREATE TABLE USING DATA_SOURCE
Property NameDefaultMeaningScope
timeZone(value of spark.sql.session.timeZone configuration)Sets the string that indicates a time zone ID to be used to format timestamps in the JSON datasources or partition values. The following formats of timeZone are supported:
  • Region-based zone ID: It should have the form ‘area/city’, such as ‘America/Los_Angeles’.
  • Zone offset: It should be in the format ‘(+|-)HH:mm’, for example ‘-08:00’ or ‘+01:00’. Also ‘UTC’ and ‘Z’ are supported as aliases of ‘+00:00’.
Other short names like ‘CST’ are not recommended to use because they can be ambiguous.
read/write
primitivesAsStringfalseInfers all primitive values as a string type.read
prefersDecimalfalseInfers all floating-point values as a decimal type. If the values do not fit in decimal, then it infers them as doubles.read
allowCommentsfalseIgnores Java/C++ style comment in JSON records.read
allowUnquotedFieldNamesfalseAllows unquoted JSON field names.read
allowSingleQuotestrueAllows single quotes in addition to double quotes.read
allowNumericLeadingZerosfalseAllows leading zeros in numbers (e.g. 00012).read
allowBackslashEscapingAnyCharacterfalseAllows accepting quoting of all character using backslash quoting mechanism.read
modePERMISSIVEAllows a mode for dealing with corrupt records during parsing.
  • PERMISSIVE: when it meets a corrupted record, puts the malformed string into a field configured by columnNameOfCorruptRecord, and sets malformed fields to null. To keep corrupt records, an user can set a string type field named columnNameOfCorruptRecord in an user-defined schema. If a schema does not have the field, it drops corrupt records during parsing. When inferring a schema, it implicitly adds a columnNameOfCorruptRecord field in an output schema.
  • DROPMALFORMED: ignores the whole corrupted records. This mode is unsupported in the JSON built-in functions.
  • FAILFAST: throws an exception when it meets corrupted records.
read
columnNameOfCorruptRecord(value of spark.sql.columnNameOfCorruptRecord configuration)Allows renaming the new field having malformed string created by PERMISSIVE mode. This overrides spark.sql.columnNameOfCorruptRecord.read
dateFormatyyyy-MM-ddSets the string that indicates a date format. Custom date formats follow the formats at datetime pattern. This applies to date type.read/write
timestampFormatyyyy-MM-dd’T’HH:mm:ss[.SSS][XXX]Sets the string that indicates a timestamp format. Custom date formats follow the formats at datetime pattern. This applies to timestamp type.read/write
timestampNTZFormatyyyy-MM-dd’T’HH:mm:ss[.SSS]Sets the string that indicates a timestamp without timezone format. Custom date formats follow the formats at Datetime Patterns. This applies to timestamp without timezone type, note that zone-offset and time-zone components are not supported when writing or reading this data type.read/write
enableDateTimeParsingFallbackEnabled if the time parser policy has legacy settings or if no custom date or timestamp pattern was provided.Allows falling back to the backward compatible (Spark 1.x and 2.0) behavior of parsing dates and timestamps if values do not match the set patterns.read
multiLinefalseParse one record, which may span multiple lines, per file. JSON built-in functions ignore this option.read
allowUnquotedControlCharsfalseAllows JSON Strings to contain unquoted control characters (ASCII characters with value less than 32, including tab and line feed characters) or not.read
encodingDetected automatically when multiLine is set to true (for reading), UTF-8 (for writing)For reading, allows to forcibly set one of standard basic or extended encoding for the JSON files. For example UTF-16BE, UTF-32LE. For writing, Specifies encoding (charset) of saved json files. JSON built-in functions ignore this option.read/write
lineSep\r, \r\n, \n (for reading), \n (for writing)Defines the line separator that should be used for parsing. JSON built-in functions ignore this option.read/write
samplingRatio1.0Defines fraction of input JSON objects used for schema inferring.read
dropFieldIfAllNullfalseWhether to ignore column of all null values or empty array during schema inference.read
localeen-USSets a locale as language tag in IETF BCP 47 format. For instance, locale is used while parsing dates and timestamps.read
allowNonNumericNumberstrueAllows JSON parser to recognize set of “Not-a-Number” (NaN) tokens as legal floating number values.
  • +INF: for positive infinity, as well as alias of +Infinity and Infinity.
  • -INF: for negative infinity, alias -Infinity.
  • NaN: for other not-a-numbers, like result of division by zero.
read
compression(none)Compression codec to use when saving to file. This can be one of the known case-insensitive shorten names (none, bzip2, gzip, lz4, snappy and deflate). JSON built-in functions ignore this option.write
ignoreNullFields(value of spark.sql.jsonGenerator.ignoreNullFields configuration)Whether to ignore null fields when generating JSON objects.write

Other generic options can be found in Generic File Source Options.