Format
Apache ORC
Summary
- Name
- Apache ORC
- Version
- V0 and V1
- Identifiers
- PUID: fmt/2030
- Format type
- Text (Structured)
- Description
- The Apache ORC (Optimized Row Columnar) format is an open-source, column-oriented data storage format which is utilized by prominent data processing frameworks including Apache Spark, Apache Hive, Apache Flink and Apache Hadoop.The format was annouced by Hortonworks (in collaboration with Facebook) in 2013.
- Note
- Specifications: https://orc.apache.org/specification/ORCv0/ https://orc.apache.org/specification/ORCv1/ Note: https://github.com/apache/orc/tree/main/examples https://en.wikipedia.org/wiki/Apache_ORC
- File extensions
-
orc - Source
- Digital Preservation Department / The National Archives
- Developed by
- The Apache Software Foundation / The Apache Software Foundation
- Supported by
- The Apache Software Foundation / The Apache Software Foundation
Internal signatures
Apache ORC
- Note
- Absolute from beginning of file, magic bytes: ORC Absolute from end of file, magic bytes: ORC.
Byte sequences
- Min Frag Length
- Absolute from BOF
- Offset
- 0
- Max offset
- 0
- Byte Sequence
4F5243- Endianness
- None
- Min Frag Length
- Absolute from EOF
- Offset
- 0
- Max offset
- 0
- Byte Sequence
4F5243(14|15|16|17|18)- Endianness
- None
Changelog
-
Added in V120
- Release date
- 25 February 2025
Apache ORC V0 and V1: Signature researched and samples provided by Digital Preservation Department, The National Archives (UK).
Apache ORC V0 and V1: Full entry added. Submitted by Digital Preservation Department, The National Archives (UK).