// Copyright 2013-2021 The Khronos Group Inc. // // SPDX-License-Identifier: CC-BY-4.0 // :regtitle: is explained in // https://discuss.asciidoctor.org/How-to-add-markup-to-author-information-in-document-title-td6488.html = glTF{tmtitle} 2.0 Specification :tmtitle: pass:q,r[^™^] :regtitle: pass:q,r[^®^] The Khronos{regtitle} 3D Formats Working Group :data-uri: :icons: font :toc2: :toclevels: 10 :sectnumlevels: 10 :max-width: 100% :numbered: :source-highlighter: coderay :title-logo-image: image:../figures/glTF_RGB_June16.svg[Logo,pdfwidth=4in,align=right] :docinfo: shared-head :stem: // This causes cross references to chapters, sections, and tables to be // rendered as "Section A.B" (for example) rather than rendering the reference // as the text of the section title. It also enables cross references to // [source] blocks as "Listing N", but only if the [source] block has a title. :xrefstyle: short :listing-caption: Listing ifndef::revdate[] :toc-placement!: [NOTE] .Note ==== Khronos posts the AsciiDoc source of the glTF specification to enable community feedback and remixing under CC-BY 4.0. Published versions of the Specification are located in the https://www.khronos.org/registry/glTF[glTF Registry]. ==== endif::[] // Table of contents is inserted here toc::[] :leveloffset: 1 [[foreword]] = Foreword Copyright 2013-2021 The Khronos Group Inc. This specification is protected by copyright laws and contains material proprietary to Khronos. Except as described by these terms, it or any components may not be reproduced, republished, distributed, transmitted, displayed, broadcast, or otherwise exploited in any manner without the express prior written permission of Khronos. This specification has been created under the Khronos Intellectual Property Rights Policy, which is Attachment A of the Khronos Group Membership Agreement available at https://www.khronos.org/files/member_agreement.pdf. Khronos grants a conditional copyright license to use and reproduce the unmodified specification for any purpose, without fee or royalty, EXCEPT no licenses to any patent, trademark or other intellectual property rights are granted under these terms. Parties desiring to implement the specification and make use of Khronos trademarks in relation to that implementation, and receive reciprocal patent license protection under the Khronos IP Policy must become Adopters under the process defined by Khronos for this specification; see https://www.khronos.org/conformance/adopters/file-format-adopter-program. Some parts of this Specification are non-normative through being explicitly identified as purely informative, and do not define requirements necessary for compliance and so are outside the Scope of this Specification. Where this Specification includes normative references to external documents, only the specifically identified sections and functionality of those external documents are in Scope. Requirements defined by external documents not created by Khronos may contain contributions from non-members of Khronos not covered by the Khronos Intellectual Property Rights Policy. Khronos makes no, and expressly disclaims any, representations or warranties, express or implied, regarding this specification, including, without limitation: merchantability, fitness for a particular purpose, non-infringement of any intellectual property, correctness, accuracy, completeness, timeliness, and reliability. Under no circumstances will Khronos, or any of its Promoters, Contributors or Members, or their respective partners, officers, directors, employees, agents or representatives be liable for any damages, whether direct, indirect, special or consequential damages for lost revenues, lost profits, or otherwise, arising from or in connection with these materials. Khronos® and Vulkan® are registered trademarks, and ANARI™, WebGL™, glTF™, NNEF™, OpenVX™, SPIR™, SPIR‑V™, SYCL™, OpenVG™ and 3D Commerce™ are trademarks of The Khronos Group Inc. OpenXR™ is a trademark owned by The Khronos Group Inc. and is registered as a trademark in China, the European Union, Japan and the United Kingdom. OpenCL™ is a trademark of Apple Inc. and OpenGL® is a registered trademark and the OpenGL ES™ and OpenGL SC™ logos are trademarks of Hewlett Packard Enterprise used under license by Khronos. ASTC is a trademark of ARM Holdings PLC. All other product names, trademarks, and/or company names are used solely for identification and belong to their respective owners. [[introduction]] = Introduction [[introduction-general]] == General This document, referred to as the "`glTF Specification`" or just the "`Specification`" hereafter, describes the glTF file format. glTF is an API-neutral runtime asset delivery format. glTF bridges the gap between 3D content creation tools and modern graphics applications by providing an efficient, extensible, interoperable format for the transmission and loading of 3D content. [[introduction-conventions]] == Document Conventions The glTF Specification is intended for use by both implementers of the asset exporters or converters (e.g., digital content creation tools) and application developers seeking to import or load glTF assets, forming a basis for interoperability between these parties. Specification text can address either party; typically, the intended audience can be inferred from context, though some sections are defined to address only one of these parties. Any requirements, prohibitions, recommendations, or options defined by <> are imposed only on the audience of that text. [[introduction-normative-terminology]] === Normative Terminology and References The key words **MUST**, **MUST NOT**, **REQUIRED**, **SHALL**, **SHALL NOT**, **SHOULD**, **SHOULD NOT**, **RECOMMENDED**, **MAY**, and **OPTIONAL** in this document are to be interpreted as described in <>. These key words are highlighted in the specification for clarity. References to external documents are considered normative if the Specification uses any of the normative terms defined in this section to refer to them or their requirements, either as a whole or in part. [[introduction-informative-language]] === Informative Language Some language in the specification is purely informative, intended to give background or suggestions to implementers or developers. If an entire chapter or section contains only informative language, its title is suffixed with "`(Informative)`". If not designated as informative, all chapters, sections, and appendices in this document are normative. All Notes, Implementation notes, and Examples are purely informative. [[introduction-technical-terminology]] === Technical Terminology The glTF Specification makes use of linear algebra terms such as *axis*, *matrix*, *vector*, etc. to identify certain math constructs and their behaviors as defined in the <>. The glTF Specification makes use of common engineering and graphics terms such as *image*, *buffer*, *texture*, etc. to identify and describe certain _glTF_ constructs and their attributes, states, and behaviors. This section defines the basic meanings of these terms in the context of the Specification. The Specification text provides fuller definitions of the terms and elaborates, extends, or clarifies the definitions. When a term defined in this section is used in normative language within the Specification, the definitions within the Specification govern and supersede any meanings the terms may have in other technical contexts (i.e. outside the Specification). accessor:: An object describing the number and the format of data elements stored in a binary buffer. animation:: An object describing the keyframe data, including timestamps, and the target property affected by it. back-facing:: See facingness. buffer:: An external or embedded resource that represents a linear array of bytes. buffer view:: An object that represents a range of a specific buffer, and optional metadata that controls how the buffer's content is interpreted. camera:: An object defining the projection parameters that are used to render a scene. facingness:: A classification of a triangle as either front-facing or back-facing, depending on the orientation (winding order) of its vertices. front-facing:: See facingness. image:: A two dimensional array of pixels encoded as a standardized bitstream, such as <>. indexed geometry:: A mesh primitive that uses a separate source of data (index values) to assemble the primitive's topology. linear blend skinning:: A skinning method that computes a per-vertex transformation matrix as a linear weighted sum of transformation matrices of the designated nodes. material:: A parametrized approximation of visual properties of the real-world object being represented by a mesh primitive. mesh:: A collection of mesh primitives. mesh primitive:: An object binding indexed or non-indexed geometry with a material. mipmap:: A set of image representations consecutively reduced by the factor of 2 in each dimension. morph target:: An altered state of a mesh primitive defined as a set of difference values for its vertex attributes. node:: An object defining the hierarchy relations and the local transform of its content. non-indexed geometry:: A mesh primitive that uses linear order of vertex attribute values to assemble the primitive's topology. normal:: A unit XYZ vector defining the perpendicular to the surface. root node:: A node that is not a child of any other node. sampler:: An object that controls how image data is sampled. scene:: An object containing a list of root nodes to render. skinning:: The process of computing and applying individual transforms for each vertex of a mesh primitive. tangent:: A unit XYZ vector defining a tangential direction on the surface. texture:: An object that combines an image and its sampler. topology type:: State that controls how vertices are assembled, e.g. as lists of triangles, strips of lines, etc. vertex attribute:: A property associated with a vertex. winding order:: The relative order in which vertices are defined within a triangle wrapping:: A process of selecting an image pixel based on normalized texture coordinates. [[introduction-normative-references]] === Normative References The following documents are referenced by normative sections of the specification: ==== External Specifications [none] * [[bcp14]] Bradner, S., _Key words for use in RFCs to Indicate Requirement Levels_, BCP 14, RFC 2119, March 1997. Leiba, B., _Ambiguity of Uppercase vs Lowercase in RFC 2119 Key Words_, BCP 14, RFC 8174, May 2017. * [[iev]] IEC 60050-102 _International Electrotechnical Vocabulary (IEV) - Part 102: Mathematics - General concepts and linear algebra_ + IEC 60050-103 _International Electrotechnical Vocabulary (IEV) - Part 103: Mathematics - Functions_ + [NOTE] .Note ==== An online version of these standards is available at https://www.electropedia.org/ ==== * [[utf8]] The Unicode Consortium, _The Unicode Standard_ * [[json]] Bray, T., Ed., _The JavaScript Object Notation (JSON) Data Interchange Format_, STD 90, RFC 8259, DOI 10.17487/RFC8259, December 2017, * [[ieee-754]] ISO/IEC 60559 _Floating-point arithmetic_ * [[png]] ISO/IEC 15948 _Portable Network Graphics (PNG): Functional specification_ + [NOTE] .Note ==== A free version of this standard is available from W3C at . ==== * [[jpeg]] ISO/IEC 10918-1 _Digital compression and coding of continuous-tone still images: Requirements and guidelines_ + [NOTE] .Note ==== This standard is equivalent to Recommendation ITU-T T.81 available at . A free version of it is available from W3C at . ==== * [[jfif]] ISO/IEC 10918-5 _Digital compression and coding of continuous-tone still images: JPEG File Interchange Format (JFIF)_ + [NOTE] .Note ==== This standard is equivalent to Recommendation ITU-T T.871 available for free at . ==== * [[exif]] CIPA DC-008-Translation-2019 _Exchangeable image file format for digital still cameras_ * [[data-uri]] Masinter, L., _The "data" URL scheme_, RFC 2397, DOI 10.17487/RFC2397, August 1998, * [[uri]] Berners-Lee, T., Fielding, R., and L. Masinter, _Uniform Resource Identifier (URI): Generic Syntax_, STD 66, RFC 3986, DOI 10.17487/RFC3986, January 2005, * [[iri]] Duerst, M. and M. Suignard, _Internationalized Resource Identifiers (IRIs)_, RFC 3987, DOI 10.17487/RFC3987, January 2005, * [[http]] Fielding, R., Ed., and J. Reschke, Ed., _Hypertext Transfer Protocol (HTTP/1.1): Message Syntax and Routing_, RFC 7230, DOI 10.17487/RFC7230, June 2014, * [[srgb]] IEC 61966-2-1 _Default RGB colour space - sRGB_ + [NOTE] .Note ==== The encoding characteristics of sRGB are freely available from ICC: https://www.color.org/chardata/rgb/srgb.xalter ==== * [[bt709]] Recommendation ITU-R BT.709-6 _Parameter values for the HDTV standards for production and international programme exchange_ * [[mikktspace]] _MikkTSpace_ * [[compositing]] Thomas Porter and Tom Duff. 1984. _Compositing digital images._ SIGGRAPH Comput. Graph. 18, 3 (July 1984), 253–259. DOI: + [NOTE] .Note ==== A free version of this paper is available from Pixar: https://graphics.pixar.com/library/Compositing/ ==== ==== Media Type Registrations [none] * [[gltf-json]] IANA. _model/gltf+json Media Type_. https://www.iana.org/assignments/media-types/model/gltf+json * [[gltf-binary]] IANA. _model/gltf-binary Media Type_. https://www.iana.org/assignments/media-types/model/gltf-binary * [[gltf-buffer]] IANA. _application/gltf-buffer Media Type_. https://www.iana.org/assignments/media-types/application/gltf-buffer * [[octet-stream]] IANA. _application/octet-stream Media Type_. https://www.iana.org/assignments/media-types/application/octet-stream * [[image-jpeg]] Freed, N. and N. Borenstein, _Multipurpose Internet Mail Extensions (MIME) Part Two: Media Types_, RFC 2046, DOI 10.17487/RFC2046, November 1996, * [[image-png]] IANA. _image/png Media Type_. https://www.iana.org/assignments/media-types/image/png [[motivation]] == Motivation and Design Goals (Informative) glTF is an open interoperable 3D asset ‘transmission’ format that is compact, and efficient to process and render at runtime. glTF 2.0 is designed to be vendor- and runtime-neutral, usable by a wide variety of native and web-based engines and applications regardless of underlying platforms and 3D graphics APIs. glTF’s focus on run-time efficiency is a different design goal than typical 3D ‘authoring’ formats. Authoring formats are typically more verbose, with higher processing overheads, to carry authoring data that is no longer needed after iterative design is complete. glTF is complementary to authoring formats, providing a common, interoperable distillation target for publishing 3D assets to a wide audience of end users. A primary goal of glTF is to be deployable on a wide range of devices and platforms, including the web and mobile devices with limited processing and memory resources. glTF can be evolved, to keep pace with growing compute capabilities over time. This helps to foster broad industry consensus on 3D functionality that can be used ubiquitously, including Physically Based Rendering. glTF combines an easily parsable JSON scene description with one or more binary resources representing geometry, animations, and other rich data. These binary resources can often be loaded directly into GPU buffers with no additional parsing or processing, combining the faithful preservation of full hierarchical scenes, nodes, meshes, cameras, materials, and animations with efficient delivery and fast loading. glTF has been designed to meet the following goals: * **Compact file sizes.** The plain text glTF JSON file description is compact and rapid to parse. All large data such as geometry, textures and animations are stored in binary files that are significantly smaller than equivalent text representations. * **Runtime-independence.** glTF is purely an asset format and does not mandate any runtime behavior. This enables its use by any application for any purpose, including display using any rendering technology, up to and including path tracing renderers. * **Complete 3D scene representation.** Not restricted to single objects, glTF can represent entire scenes, including nodes, transformations, transform hierarchy, meshes, materials, cameras, and animations. * **Extensibility.** glTF is fully extensible, enabling the addition of both general-purpose and vendor-specific extensions, including geometry and texture compression. Widely adopted extensions may be considered for integration into future versions of the glTF specification. The following are outside the scope of glTF 2.0: * **glTF is not a streaming format.** The binary data in glTF is inherently streamable, and the buffer design allows for fetching data incrementally, but there are no other streaming constructs in glTF 2.0. * **glTF is not an authoring format.** glTF deliberately does not retain 3D authoring information, in order to preserve runtime efficiency, however glTF files may be ingested by 3D authoring tools for remixing. * **glTF is not intended to be human-readable,** though by virtue of being represented in JSON, it is developer-friendly. [[gltf-basics]] == glTF Basics A glTF asset is represented by: * A JSON-formatted file (`.gltf`) containing a full scene description: node hierarchy, materials, cameras, as well as descriptor information for meshes, animations, and other constructs. * Binary files (`.bin`) containing geometry, animation, and other buffer-based data. * Image files (`.jpg`, `.png`) containing texture images. Binary and image resources **MAY** also be embedded directly in JSON using <> or stored side-by-side with JSON in <> container. A valid glTF asset **MUST** specify its version. [[versioning]] == Versioning Any updates made to the glTF Specification in a minor version **MUST** be backward and forward compatible. Backward compatibility means that any client implementation that supports loading a glTF 2.x asset will also be able to load a glTF 2.0 asset. Forward compatibility means that a client implementation that only supports glTF 2.0 can load glTF 2.x assets while gracefully ignoring any new features it does not understand. A minor version update **MAY** introduce new features but **MUST NOT** change any previously existing behavior. Existing functionality **MAY** be deprecated in a minor version update, but it **MUST NOT** be removed. Major version updates **MAY** be incompatible with previous versions. [[file-extensions-and-media-types]] == File Extensions and Media Types * <> glTF files **SHOULD** use `.gltf` extension and <> Media Type. * glTF files stored in <> container **SHOULD** use `.glb` extension and <> Media Type. * Files representing binary buffers **SHOULD** use either: ** `.bin` file extension with <> Media Type; ** `.bin`, `.glbin`, or `.glbuf` file extensions with <> Media Type. [[json-encoding]] == JSON Encoding Although glTF Specification does not define any subset of the <> format, implementations **SHOULD** be aware of its peculiar properties that could affect asset interoperability. 1. glTF JSON data **SHOULD** be written with UTF-8 encoding without BOM. This requirement is not applied when a glTF implementation does not control string encoding. glTF implementations **SHOULD** adhere to <>, Section 8.1. with regards to treating BOM presence. 2. ASCII characters stored in glTF JSON **SHOULD** be written without JSON escaping. + [NOTE] .Example ==== `"buffer"` instead of `"\u0062\u0075\u0066\u0066\u0065\u0072"`. ==== 3. Non-ASCII characters stored in glTF JSON **MAY** be escaped. + [NOTE] .Example ==== These two examples represent the same glTF JSON data. [source,json] ----- { "asset": { "version": "2.0" }, "nodes": [ { "name": "куб" }, { "name": "立方體" } ] } ----- [source,json] ----- { "asset": { "version": "2.0" }, "nodes": [ { "name": "\u043a\u0443\u0431" }, { "name": "\u7acb\u65b9\u9ad4" } ] } ----- ==== 4. Property names (keys) within JSON objects **SHOULD** be unique. glTF client implementations **SHOULD** override lexically preceding values for the same key. 5. Some of glTF properties are defined as integers in the schema. Such values **MAY** be stored as decimals with a zero fractional part or by using exponent notation. Regardless of encoding, such properties **MUST NOT** contain any non-zero fractional value. + [NOTE] .Example ==== `100`, `100.0`, and `1e2` represent the same value. See <>, Section 6 for more details. ==== 6. Non-integer numbers **SHOULD** be written in a way that preserves original values when these numbers are read back, i.e., they **SHOULD NOT** be altered by JSON serialization / deserialization roundtrip. + [NOTE] .Implementation Note ==== This is typically achieved with algorithms like Grisu2 used by common JSON libraries. ==== [[uris]] == URIs glTF assets use <> or <> to reference buffers and image resources. Assets **MAY** contain at least these two URI types: - **Data URIs** that embed binary resources in the glTF JSON as defined by the <>. The Data URI's `mediatype` field **MUST** match the encoded content. + [NOTE] .Implementation Note ==== Base64 encoding used in Data URI increases the payload's byte length by 33%. ==== - **Relative paths** -- `path-noscheme` or `ipath-noscheme` as defined by <>, Section 4.2 or <>, Section 2.2 -- without scheme, authority, or parameters. Reserved characters (as defined by <>, Section 2.2. and <>, Section 2.2.) **MUST** be percent-encoded. Paths with non-ASCII characters **MAY** be written as-is, with JSON string escaping, or with percent-encoding; all these options are valid. For example, the following three paths point to the same resource: [source,json] ---- { "images": [ { "uri": "grande_sphère.png" }, { "uri": "grande_sph\u00E8re.png" }, { "uri": "grande_sph%C3%A8re.png" } ] } ---- Client implementations **MAY** optionally support additional URI components. For example `http://` or `file://` schemes, authorities, hostnames, absolute paths, and query or fragment parameters. Assets containing these additional URI components would be less portable. [NOTE] .Implementation Note ==== This allows the application to decide the best approach for delivery: if different assets share many of the same geometries, animations, or textures, separate files may be preferred to reduce the total amount of data requested. With separate files, applications may progressively load data and do not need to load data for parts of a model that are not visible. If an application cares more about single-file deployment, embedding data may be preferred even though it increases the overall size due to base64 encoding and does not support progressive or on-demand loading. Alternatively, an asset could use the GLB container to store JSON and binary data in one file without base64 encoding. See <> for details. ==== URIs **SHOULD** undergo syntax-based normalization as defined by <>, Section 6.2.2, <>, Section 5.3.2, and applicable schema rules (e.g., <>, Section 2.7.3 for HTTP) on export and/or import. [NOTE] .Implementation Note ==== While the specification does not explicitly disallow non-normalized URIs, their use may be unsupported or lead to unwanted side-effects -- such as security warnings or cache misses -- on some platforms. ==== [[concepts]] = Concepts [[concepts-general]] == General The figure below shows relations between top-level arrays in a glTF asset. See the <>. .glTF Object Hierarchy image::figures/objects.svg[pdfwidth=6in,align=left] [[asset]] == Asset Each glTF asset **MUST** have an `asset` property. The `asset` object **MUST** contain a `version` property that specifies the target glTF version of the asset. Additionally, an optional `minVersion` property **MAY** be used to specify the minimum glTF version support required to load the asset. Both `version` and `minVersion` properties **MUST** conform to the asset version string syntax defined as follows. The version string is a concatenation of a major version component, a single "`dot`" character (`0x2E`), and a minor version component. Each version component is either a string consisting of a single decimal "`zero`" character (`0x30`), or a single non-zero decimal character (with codes from `0x31` to `0x39`) optionally followed by up to eight any decimal characters (with codes from `0x30` to `0x39`). [NOTE] ==== In other words, a version string is two decimal numbers from 0 to 999999999 without leading zeros joined with a single dot. ==== [NOTE] .Example valid version strings ==== ````` "0.0", "0.12", "2.0", "12.34", "999999999.999999999" ````` ==== [NOTE] .Example invalid version strings ==== ````` "+1.0", "2.0.1", "2.00", "01.1" ````` ==== A version `"A.B"` is less than a version `"C.D"`, if and only if `A` is less than `C`, or if `A` is equal to `C` and `B` is less than `D`, where `A`, `B`, `C`, and `D` are version components of the version strings as described above. [NOTE] .Version string comparison examples ==== - `"2.0"` is less than `"3.0"` - `"2.0"` is less than `"2.1"` - `"2.1"` is less than `"2.10"` - `"2.10"` is less than `"3.0"` ==== The `minVersion` property allows asset creators to specify a minimum version that a client implementation **MUST** support in order to load the asset. If present, the minimum version specified by the `minVersion` property **MUST** be less than or equal to the version specified by the `version` property. [NOTE] .Implementation Note ==== Client implementations should first check whether a `minVersion` property is specified and ensure both its major and minor versions can be supported. If no `minVersion` is specified, then clients should check the `version` property and ensure its major version is supported. ==== [CAUTION] .Implementation Note ==== The version specified in the <> header refers only to the GLB container version and is not related to the version of glTF JSON. ==== Additional metadata **MAY** be stored in optional properties such as `generator` or `copyright`. For example, [source,json] ---- { "asset": { "version": "2.0", "generator": "collada2gltf@f356b99aef8868f74877c7ca545f2cd206b9d3b7", "copyright": "2017 (c) Khronos Group" } } ---- [[indices-and-names]] == Indices and Names Entities of a glTF asset are referenced by their indices in corresponding arrays, e.g., a `bufferView` refers to a `buffer` by specifying the buffer's index in `buffers` array. For example: [source,json] ---- { "buffers": [ { "byteLength": 1024, "uri": "path-to.bin" } ], "bufferViews": [ { "buffer": 0, "byteLength": 512, "byteOffset": 0 } ] } ---- In this example, `buffers` and `bufferViews` arrays have only one element each. The bufferView refers to the buffer using the buffer's index: `"buffer": 0`. Indices **MUST** be non-negative integer numbers. Indices **MUST** always point to existing elements. Whereas indices are used for internal glTF references, optional _names_ are used for application-specific uses such as display. Any top-level glTF object **MAY** have a `name` string property for this purpose. These property values are not guaranteed to be unique as they are intended to contain values created when the asset was authored. For property names, glTF usually uses camel case, `likeThis`. [[coordinate-system-and-units]] == Coordinate System and Units glTF uses a right-handed coordinate system. glTF defines +Y as up; the front side of a glTF asset faces +Z, the left side of a glTF asset faces +X. .glTF Coordinate System Orientation image::figures/coordinate-system.svg[pdfwidth=2in,align=left] The units for all linear distances are meters. All angles are in radians. Red, Green, and Blue primary colors use <> chromaticity coordinates. [NOTE] .Implementation Note ==== Chromaticity coordinates define the interpretation of each primary color channel of the color model. In the context of a typical display, color primaries describe the color of the red, green, and blue phosphors or filters. Unless a wide color gamut output is explicitly used, client implementations usually do not need to convert colors. Future specification versions or extensions may allow other color primaries (such as P3). ==== [[scenes]] == Scenes [[scenes-overview]] === Overview glTF 2.0 assets **MAY** contain zero or more _scenes_, the set of visual objects to render. Scenes are defined in a `scenes` array. All nodes listed in `scene.nodes` array **MUST** be root nodes, i.e., they **MUST NOT** be listed in a `node.children` array of any node. The same root node **MAY** appear in multiple scenes. An additional root-level property, `scene` (note singular), identifies which of the scenes in the array **SHOULD** be displayed at load time. When `scene` is undefined, client implementations **MAY** delay rendering until a particular scene is requested. A glTF asset that does not contain any scenes **SHOULD** be treated as a library of individual entities such as materials or meshes. The following example defines a glTF asset with a single scene that contains a single node. [source,json] ---- { "nodes": [ { "name": "singleNode" } ], "scenes": [ { "name": "singleScene", "nodes": [ 0 ] } ], "scene": 0 } ---- [[nodes-and-hierarchy]] === Nodes and Hierarchy glTF assets **MAY** define _nodes_, that is, the objects comprising the scene to render. Nodes **MAY** have transform properties, as described later. Nodes are organized in a parent-child hierarchy known informally as the _node hierarchy_. A node is called a _root node_ when it doesn't have a parent. The node hierarchy **MUST** be a set of disjoint strict trees. That is node hierarchy **MUST NOT** contain cycles and each node **MUST** have zero or one parent node. A node lists its children node indices in the `children` array property. Each of those nodes could in turn have its own children, creating a hierarchy of nodes. [NOTE] .Example ==== In the following example, the node named `Car` has four children each of which has one child node. [source,json] ---- { "nodes": [ { "name": "Car", "children": [1, 2, 3, 4] }, { "name": "wheel_1", "children": [5] }, { "name": "wheel_2", "children": [6] }, { "name": "wheel_3", "children": [7] }, { "name": "wheel_4", "children": [8] }, { "name": "tire_1" }, { "name": "tire_2" }, { "name": "tire_3" }, { "name": "tire_4" } ] } ---- ==== [[transformations]] === Transformations By default, a node's local space transform is identity. Any node **MAY** define a local space transform either by supplying a precomputed transformation matrix in the `matrix` property, or any of `translation`, `rotation`, and `scale` properties (also known as _TRS properties_). _Translation_ and _scale_ are 3D vectors in the local coordinate system; _rotation_ is a quaternion value stored in XYZW order where W is the scalar. If defined, the rotation quaternion **MUST** be unit. When the `matrix` property is not defined, the local transform is computed from the TRS properties as follows, where - stem:[t_x], stem:[t_y], and stem:[t_z] are the translation vector components; - stem:[r_x], stem:[r_y], stem:[r_z], and stem:[r_w] are the rotation quaternion components; - stem:[s_x], stem:[s_y], and stem:[s_z] are the scale vector components. [stem] +++++ ((s_x * (1 - 2(r_y^2 + r_z^2)), s_y * 2(r_xr_y - r_zr_w), s_z * 2(r_xr_z + r_yr_w), t_x), (s_x * 2(r_xr_y + r_zr_w), s_y * (1 - 2(r_x^2 + r_z^2)), s_z * 2(r_yr_z - r_xr_w), t_y), (s_x * 2(r_xr_z - r_yr_w), s_y * 2(r_yr_z + r_xr_w), s_z * (1 - 2(r_x^2 + r_y^2)), t_z), (0, 0, 0, 1)) +++++ When the `matrix` property is defined, it **MUST** be decomposable to TRS properties, i.e., the specified matrix **MUST** conform to the following conditions: - the last row is stem:[(0, 0, 0, 1)]; - none of the first three columns is all-zeros; - the absolute value of the determinant of the upper-left 3x3 matrix formed by dividing each matrix element by the length of its column is close to positive one. [NOTE] .Implementation Note ==== These conditions ensure that the matrix has the following properties: - the matrix does not contain shear or perspective transforms; - decomposing the matrix into TRS properties and composing them back would produce nearly the same matrix, subject to floating-point precision. The `matrix` property cannot represent zero scale; assets would need to use the explicit `scale` property for that. ==== The global transformation matrix of a node is the product of the global transformation matrix of its parent node and its own local transformation matrix. When the node has no parent node, its global transformation matrix is identical to its local transformation matrix. [NOTE] .Example ==== In the example below, a node named `Box` defines non-default rotation and translation. [source,json] ---- { "nodes": [ { "name": "Box", "rotation": [ 0, 0, 0, 1 ], "scale": [ 1, 1, 1 ], "translation": [ -17.7082, -11.4156, 2.0922 ] } ] } ---- The next example defines the transform for a node with attached camera using the `matrix` property rather than using the individual TRS values: [source,json] ---- { "nodes": [ { "name": "node-camera", "camera": 1, "matrix": [ -0.99975, -0.00679829, 0.0213218, 0, 0.00167596, 0.927325, 0.374254, 0, -0.0223165, 0.374196, -0.927081, 0, -0.0115543, 0.194711, -0.478297, 1 ] } ] } ---- ==== Computing local and global transformation matrices **SHOULD** be done in double precision. [NOTE] .Implementation Note ==== Keeping double precision for intermediate values usually helps with preventing visual artifacts that could be caused by specific combinations of vertex data and/or node transforms. Since most GPUs do not directly support double-precision data, final values might still be truncated to single precision prior to rendering. ==== [NOTE] .Implementation Note ==== When the scale is zero on all three axes (by node transform or by animated scale), implementations are free to optimize away rendering of the node's mesh, and all of the node's children's meshes. This provides a mechanism to animate visibility. Skinned meshes must not use this optimization unless all of the joints in the skin are scaled to zero simultaneously. ==== [[binary-data-storage]] == Binary Data Storage [[buffers]] === Buffers [[buffers-overview]] ==== Overview A _buffer_ is arbitrary data stored as a binary blob. Buffers **MAY** contain any combination of any data used by the glTF asset. Binary blobs allow efficient creation of GPU buffers and textures since they require no additional parsing, except perhaps decompression. A buffer is defined by the `byteLength` and `uri` properties that specify the byte size of the buffer and the URI to the buffer data, respectively. glTF assets **MAY** have any number of buffers. Buffers are defined in the asset's `buffers` array. While there's no hard upper limit on a buffer's byte size, glTF assets **SHOULD NOT** use buffers bigger than 2^53^-1 bytes because some JSON parsers may be unable to parse the value of their `byteLength` property correctly. Buffers stored as <> binary chunk have an implicit size limit of 2^32^-1 bytes. [NOTE] .Example ==== [source,json] ---- { "buffers": [ { "byteLength": 102040, "uri": "duck.bin" } ] } ---- ==== The byte length of the referenced resource **MUST** be greater than or equal to the buffer's `byteLength` property. Only the resource's byte range from zero (inclusive) to `byteLength` (exclusive) is considered referenced by the buffer object. Buffer data **MAY** alternatively be embedded in the glTF file via `data:` URIs. When a `data:` URI is used for buffer storage, its `mediatype` field **MUST** be set to `application/gltf-buffer` or `application/octet-stream`. If the `uri` property is undefined, the buffer's data is provided by other means, such as a GLBv2 chunk (see below) or additional extensions. In absence of data sources, the buffer's content is undefined and handling of such buffers is implementation-dependent. [[glb-stored-buffer]] ==== GLB-stored Buffer The glTF asset **MAY** use the GLBv2 file container format to pack glTF JSON and one glTF buffer into the same file. Data for that buffer is provided by the GLB-stored `BIN` chunk. When the glTF asset is stored inside a GLBv2 file, a glTF buffer that has index zero, i.e., the first element of the `buffers` array, and has its `uri` property undefined represents the GLB-stored buffer. The GLB file **MUST** have a `BIN` chunk in this case. [NOTE] .Example ==== In the following snippet, the first buffer object refers to GLB-stored data, while the second buffer uses an external resource. [source,json] ---- { "buffers": [ { "byteLength": 35884 }, { "byteLength": 504, "uri": "external.bin" } ] } ---- ==== Any glTF buffer with undefined `uri` property that is not the first element of the `buffers` array does not refer to the GLB-stored BIN chunk, and the behavior of such buffers is left undefined to accommodate future extensions and specification versions. The byte length of the GLB's `BIN` chunk **MUST** be greater than or equal to the byte length of the corresponding glTF buffer. [NOTE] .Implementation Note ==== Not requiring strict equality of chunk's and buffer's lengths is consistent with how buffers use external resources and slightly simplifies glTF to GLBv2 conversion: buffer's `byteLength` does not need to be updated after applying GLBv2 padding. ==== See <> for details on GLBv2 File Format. [[buffer-views]] === Buffer Views [[buffer-views-overview]] ==== Overview A _buffer view_ represents a span of bytes defined by the `byteLength`, `buffer`, `byteOffset`, `byteStride`, and `target` properties. The `byteLength` property defines the buffer view's byte length, the `buffer` and `byteOffset` properties define the buffer to use as the data source and the byte offset within it, respectively, the `byteStride` property defines the stride when the buffer view is used for vertex attributes, and the optional `target` property hints at the buffer usage when transferring mesh data. Buffer views are defined in the asset's `bufferViews` array. If the `byteOffset` property is not defined, it is assumed to be zero. [NOTE] .Example ==== The following snippet defines two buffer views: the first holds the vertex indices for an indexed mesh primitive, and the second holds the vertex data for a mesh primitive. [source,json] ---- { "bufferViews": [ { "buffer": 0, "byteLength": 25272, "byteOffset": 0, "target": 34963 }, { "buffer": 0, "byteLength": 76768, "byteOffset": 25272, "byteStride": 32, "target": 34962 } ] } ---- ==== The referenced buffer **MUST** have enough bytes for the buffer view, i.e., the buffer's byte length **MUST** be greater than or equal to the sum of the buffer view's `byteOffset` and `byteLength` property values. Buffer views directly used by vertex indices accessors, vertex attribute accessors, or inverse bind matrices accessors **MUST NOT** contain more than one kind of data, e.g., if a buffer view is used by a vertex attribute accessor, that buffer view cannot be used by anything else except other vertex attribute accessors. Buffer views used by images or any other objects introduced by extensions or future Specification versions that refer to a buffer view as a whole, i.e., without specifying byte offset and/or byte length, **MUST NOT** be used by any objects that can refer to buffer view regions, such as accessors. When a buffer view is directly used by vertex indices accessors or vertex attribute accessors, it **MAY** define the `target` property with a value of _element array buffer_ or _array buffer_, respectively. The `target` value uses integer enums defined in the <<_bufferview_target,Properties Reference>>. Buffer views used for other kinds of data **MUST NOT** define the `target` property. [NOTE] .Implementation Note ==== This allows client implementations to early designate each buffer view to a proper processing step, e.g., data from buffer views with vertex indices and vertex attributes would be copied to the appropriate GPU buffers, while buffer views with image data would be passed to format-specific image decoders. ==== When a buffer view is used for vertex attribute data, it **MAY** define the `byteStride` property. This property specifies the stride in bytes between the first bytes of any two consecutive vertex attribute values. If two or more vertex attributes use the same buffer view, that buffer view **MUST** define its byte stride. Buffer views with other kinds of data **MUST NOT** define the byte stride. If the `byteStride` property is defined, all of the following restrictions apply to it: - its value **MUST** be a multiple of four; - its value **MUST** be greater than or equal to 4 and less than or equal to 252; - its value **MUST** be less than or equal to the value of the `byteLength` property. [[accessors]] === Accessors [[accessors-overview]] ==== Overview Buffers and buffer views do not contain type information. They simply define the byte ranges for retrieval from the referenced resources. Structured binary data needed by objects within the glTF asset, such as meshes, skins, and animations, is accessed via _accessors_. Accessors are stored in the asset's `accessors` array. An accessor provides a finite non-empty sequence of typed values. The number of elements provided by the accessor is defined by its `count` property. The dimensionality of accessor elements is defined by the `type` property and the data type of those elements is defined by the `componentType` and `normalized` properties. If the `bufferView` accessor property is defined, the accessor elements are sourced from the referenced buffer view and the `byteOffset` property defines the byte offset of the first accessed element within it. If the `byteOffset` property is not defined, it is assumed to be zero. If the buffer view's `byteStride` property is present, it defines the stride in bytes between the first bytes of consecutive accessor elements. If the `bufferView` accessor property is not defined, the accessor elements are not sourced from a buffer view and the `byteOffset` property **MUST NOT** be defined. An accessor element source **MAY** be defined by an extension. In absence of element sources, accessor elements are sourced from an infinite sequence of zero bytes. If the `sparse` accessor property is defined, the accessor elements are selectively replaced based on the properties of the `sparse` object. Such replacements happen after resolving the element source. Optional `min` and `max` properties provide minimum and maximum values that the accessor can provide, respectively. [NOTE] .Example ==== The following snippet shows two accessors, the first is a scalar 16-bit integer accessor, and the second is a floating-point three-component vector accessor. [source,json] ---- { "accessors": [ { "bufferView": 0, "byteOffset": 0, "componentType": 5123, "count": 12636, "type": "SCALAR" }, { "bufferView": 1, "byteOffset": 12, "componentType": 5126, "count": 2399, "max": [ 0.9617, 1.6397, 0.5392 ], "min": [ -0.6929, 0.0992, -0.6132 ], "type": "VEC3" } ] } ---- ==== The following sections provide more details on each aspect of accessors. [[accessor-data-types]] ==== Accessor Data Types The effective component type of accessor elements is defined by the combination of the enumerated `componentType` and boolean `normalized` properties. The following table lists effective component types with their corresponding accessor properties and short names used in the subsequent sections of the Specification. [options="header",cols="10%,30%,10%,10%,10%"] |==== | Short Name | Description | Size in Bytes | `componentType` | `normalized` | _**sint8**_ | 8-bit signed integer | 1 | `5120` | false | _**snorm8**_ | 8-bit signed normalized integer | 1 | `5120` | true | _**uint8**_ | 8-bit unsigned integer | 1 | `5121` | false | _**unorm8**_ | 8-bit unsigned normalized integer | 1 | `5121` | true | _**sint16**_ | 16-bit signed integer | 2 | `5122` | false | _**snorm16**_ | 16-bit signed normalized integer | 2 | `5122` | true | _**uint16**_ | 16-bit unsigned integer | 2 | `5123` | false | _**unorm16**_ | 16-bit unsigned normalized integer | 2 | `5123` | true | _**uint32**_ | 32-bit unsigned integer | 4 | `5125` | false | _**float32**_ | 32-bit signed <> float | 4 | `5126` | false |==== All multi-byte component types use little endian byte order. Signed integers use two's complement representation. The `normalized` property **MUST NOT** be true when the `componentType` is `5125` or `5126`. Only the values of the `componentType` property present in the table above are in scope of this Specification; support for others **MAY** be added by extensions or future Specification versions. Unsigned normalized integers represent floating-point numbers in the range latexmath:[[0, 1\]]. Signed normalized integers represent floating-point numbers in the range latexmath:[[-1, +1\]]. The conversions for the supported `componentType` property values are defined as follows, where stem:[c] is the accessed integer value and stem:[f] is the effective floating-point value. These conversions **SHOULD** use at least 32-bit floating-point precision. [options="header",cols="15%,20%"] |==== | Short Name | Conversion | _**snorm8**_ | latexmath:[f = \frac{\max(c, -127)}{127.0}] | _**unorm8**_ | latexmath:[f = \frac{c}{255.0}] | _**snorm16**_ | latexmath:[f = \frac{\max(c, -32767)}{32767.0}] | _**unorm16**_ | latexmath:[f = \frac{c}{65535.0}] |==== [NOTE] .Implementation Note ==== Signed normalized representations have redundant encodings. Both `-127` and `-128` represent `-1.0` for _snorm8_; both `-32767` and `-32768` represent `-1.0` for _snorm16_. ==== [NOTE] .Implementation Note ==== When accessor elements are used for vertex attributes, these normalization conversions are usually performed by GPU hardware. ==== [[accessor-components]] ==== Accessor Components The number and semantics of components per a single accessor element are defined by the `type` property as follows. [options="header",cols="5%,10%,10%"] |==== | `type` | Description | Number of components | `"SCALAR"`| Scalar value | 1 | `"VEC2"` | Two-component vector | 2 | `"VEC3"` | Three-component vector | 3 | `"VEC4"` | Four-component vector | 4 | `"MAT2"` | 2x2 matrix | 4 | `"MAT3"` | 3x3 matrix | 9 | `"MAT4"` | 4x4 matrix | 16 |==== The vector components are stored in `(x, y, z, w)` order. The matrix elements are stored in the column-major order. [[accessor-element-size]] ==== Accessor Element Size The byte size of one non-matrix accessor element is a product of its component type byte size and the number of components as defined above. The following table provides element byte sizes for all supported combinations of non-matrix accessor types and component byte sizes. [options="header",cols="5%,10%,10%"] |==== | `type` | Component size in bytes | Element size in bytes | `"SCALAR"` | 1 | 1 | `"SCALAR"` | 2 | 2 | `"SCALAR"` | 4 | 4 | `"VEC2"` | 1 | 2 | `"VEC2"` | 2 | 4 | `"VEC2"` | 4 | 8 | `"VEC3"` | 1 | 3 | `"VEC3"` | 2 | 6 | `"VEC3"` | 4 | 12 | `"VEC4"` | 1 | 4 | `"VEC4"` | 2 | 8 | `"VEC4"` | 4 | 16 |==== The byte size of one matrix accessor element is a product of the number of matrix columns and the byte size of one matrix column. The byte size of one matrix column is the accessor component type byte size multiplied by the number of matrix rows and rounded up to the nearest multiple of four. The following table provides element byte sizes for all supported combinations of matrix accessor types and component byte sizes. [options="header",cols="5%,10%,10%"] |==== | `type` | Component size in bytes | Element size in bytes | `"MAT2"` | 1 | 8 | `"MAT2"` | 2 | 8 | `"MAT2"` | 4 | 16 | `"MAT3"` | 1 | 12 | `"MAT3"` | 2 | 24 | `"MAT3"` | 4 | 36 | `"MAT4"` | 1 | 16 | `"MAT4"` | 2 | 32 | `"MAT4"` | 4 | 64 |==== Data layouts for accessors of `"MAT2"` type and any one-byte component type as well as accessors of `"MAT3"` type and any one- or two-byte component types include extra padding bytes (marked as `X`) as follows. .Matrix 2x2, 1-byte components image::figures/padding-mat2-1byte.svg[pdfwidth=2in,align=left] .Matrix 3x3, 1-byte components image::figures/padding-mat3-1byte.svg[pdfwidth=2in,align=left] .Matrix 3x3, 2-byte components image::figures/padding-mat3-2byte.svg[pdfwidth=4in,align=left] [[data-alignment]] ==== Accessing Data from Buffer Views When the accessor's `bufferView` property is defined, the referenced buffer view is the source of accessor elements. Let: - _offset_ be the value of the accessor's `byteOffset` property; - _size_ be the accessor's element size in bytes as defined above; - _stride_ be the value of the buffer view's `byteStride` property if it is defined or _size_ otherwise; - _count_ be the value of the accessor's `count` property. Then the byte offset of the _i_-th accessor element within the buffer view is defined by the following expression where _i_ is the accessor element index from zero (inclusive) to _count_ (exclusive). [latexmath] +++++ \mathit{offset} + i \times \mathit{stride} +++++ [NOTE] .Implementation Note ==== When _stride_ is greater than _size_, accessor elements are not tightly-packed and thus interleaved bytes would need to be skipped during iteration over accessor elements. Graphics APIs usually support that directly by accepting byte stride values associated with vertex attributes. ==== The referenced buffer view **MUST** be large enough to fit all accessed elements, i.e., the buffer view's byte length **MUST** be greater than or equal to the value of the following expression. [latexmath] +++++ \mathit{offset} + (\mathit{count} - 1) \times \mathit{stride} + \mathit{size} +++++ The accessor's byte offset into the buffer view and the buffer view's byte offset into the buffer **MUST** be multiples of the accessor's component type byte size. [NOTE] .Implementation Note ==== These alignment requirements allow client implementations to more efficiently process binary buffers because creating aligned data views usually does not require extra copying. ==== Accessors used for vertex attributes **MAY** reference buffer views with a defined `byteStride` property; accessors that are not used for vertex attributes **MUST NOT** reference buffer views that define byte stride. If two or more vertex attribute accessors use the same buffer view, that buffer view **MUST** define its byte stride. If the referenced buffer view defines its byte stride, the following restrictions apply in addition to those mentioned in the section about buffer views: - The byte stride of the referenced buffer view **MUST** be greater than or equal to the accessor's element byte size. + [NOTE] .Rationale ==== This restriction ensures that adjacent accessor elements do not overlap. ==== - The byte stride of the referenced buffer view **MUST** be a multiple of the accessor's component type byte size. + [NOTE] .Rationale ==== This restriction is currently redundant because the byte stride value has to be a multiple of four anyway and that implicitly covers all component type byte sizes defined in this Specification. However, if an extension or a future Specification version add support for accessors with another component type byte size, e.g., greater than four, the byte stride of buffer views used with such accessors would have to be a multiple of that size. ==== If the accessor is used for vertex attributes, its elements **MUST** be aligned to 4-byte boundaries inside the buffer view, i.e., the accessor's byte offset **MUST** be a multiple of four and either the component type byte size **MUST** be a multiple of four or the buffer view's byte stride **MUST** be defined. [NOTE] .Implementation Note ==== Multiple GPU platforms require vertex data alignment for performance and/or compatibility reasons. ==== [NOTE] .Example ==== The following snippet defines two accessors that refer to a buffer view with a byte stride greater than the byte size of the accessor elements. The first accessor provides 3 two-component vectors of the _unorm16_ component type and the second accessor provides 3 three-component vectors of the _unorm8_ type. The first accessor retrieves data from the `[0, 3]`, `[8, 11]`, and `[16, 19]` inclusive byte ranges of the buffer view, and the second accessor retrieves data from the `[4, 6]`, `[12, 14]`, and `[20, 22]` inclusive byte ranges of the buffer view. Note that bytes `7` and `15` are not accessed, and the byte `23` does not even have to be included in the buffer view (as per the rules of this section). [source,json] ---- { "bufferViews": [ { "buffer": 0, "byteLength": 23, "byteStride": 8 } ], "accessors": [ { "bufferView": 0, "componentType": 5123, "count": 3, "normalized": true, "type": "VEC2" }, { "bufferView": 0, "byteOffset": 4, "componentType": 5121, "count": 3, "normalized": true, "type": "VEC3" } ] } ---- ==== [[sparse-accessors]] ==== Sparse Accessors [[sparse-accessors-overview]] ===== Overview Sparse data encoding is usually more memory-efficient than dense encoding when describing incremental changes with respect to reference values. This is often the case when encoding morph targets; it is, in general, more efficient to describe a few displaced vertices in a morph target than transmitting all morph target vertices. The `sparse` property of the accessor object defines which accessor elements are replaced. The `sparse` property value is a JSON object that contains the following **REQUIRED** properties: - `count`: the number of replaced elements, it **MUST** be greater than zero and less than or equal to the value of the accessor's `count` property; - `indices`: the object describing the location and the component type of indices of accessor elements to be replaced; - `values`: the object describing the location of the replacement elements corresponding to the indices. [NOTE] .Example ==== The following snippet shows a sparse accessor with ten replaced elements. [source,json] ---- { "accessors": [ { "bufferView": 0, "byteOffset": 0, "componentType": 5126, "count": 12636, "type": "VEC3", "sparse": { "count": 10, "indices": { "bufferView": 1, "byteOffset": 0, "componentType": 5123 }, "values": { "bufferView": 2, "byteOffset": 0 } } } ] } ---- ==== [[sparse-accessor-indices]] ===== Sparse Accessor Indices The `indices` object contains three properties: `componentType`, `bufferView`, and `byteOffset`. The `componentType` property defines the component type of the sparse index values and it **MUST** correspond to any of the unsigned integer data types defined above. The `bufferView` and `byteOffset` properties define the buffer view containing the indices of the replaced elements and the byte offset of the first index value within the buffer view, respectively. If the `byteOffset` property is not defined, it is assumed to be zero. The referenced buffer view **MUST NOT** have its `byteStride` or `target` properties defined. The index values are tightly packed, i.e., the distance between the first bytes of adjacent index values is equal to the byte size of the index component type. Let: - _sparseCount_ be the number of replaced elements; - _sparseIndexSize_ be the byte size of the sparse index component type; - _sparseIndicesOffset_ be the value of the `byteOffset` property of the `indices` object. The _sparseIndicesOffset_ and the buffer view's byte offset into the buffer **MUST** be multiples of the _sparseIndexSize_. Then the byte offset of the _j_-th sparse index within the sparse indices buffer view is defined by the following expression where _j_ is the sparse index from zero (inclusive) to _sparseCount_ (exclusive). [latexmath] +++++ \mathit{sparseIndicesOffset} + j \times \mathit{sparseIndexSize} +++++ The referenced buffer view **MUST** be large enough to fit all indices, i.e., the byte length of the referenced buffer view **MUST** be greater than or equal to the value of the following expression. [latexmath] +++++ \mathit{sparseIndicesOffset} + \mathit{sparseCount} \times \mathit{sparseIndexSize} +++++ The sequence of stored indices **MUST** be strictly increasing and all index values **MUST** be less than the value of the accessor's `count` property. [[sparse-accessor-values]] ===== Sparse Accessor Values The `values` object contains `bufferView` and `byteOffset` properties. They define the buffer view containing the replacement accessor elements and the byte offset of the first replacement element within the buffer view, respectively. If the `byteOffset` property is not defined, it is assumed to be zero. The referenced buffer view **MUST NOT** have its `byteStride` or `target` properties defined. The replacement elements have the same dimensionality and component type as the original accessor elements. The replacement elements are always tightly packed, i.e., the distance between the first bytes of adjacent replacement elements is equal to the accessor element byte size. Let: - _sparseCount_ be the number of replaced elements; - _size_ be the accessor's element size in bytes; - _sparseValuesOffset_ be the value of the `byteOffset` property of the `values` object. The _sparseValuesOffset_ and the buffer view's byte offset into the buffer **MUST** be multiples of the accessor's component byte size. Then the byte offset of the _j_-th replacement element within the sparse values buffer view is defined by the following expression where _j_ is the replacement element index from zero (inclusive) to _sparseCount_ (exclusive). [latexmath] +++++ \mathit{sparseValuesOffset} + j \times \mathit{size} +++++ The referenced buffer view **MUST** be large enough to fit all replacement elements, i.e., the byte length of the referenced buffer view **MUST** be greater than or equal to the value of the following expression. [latexmath] +++++ \mathit{sparseValuesOffset} + \mathit{sparseCount} \times \mathit{size} +++++ [[sparse-accessor-usage]] ===== Sparse Accessor Usage When the accessor has sparse data, retrieving the _i_-th accessor element is done by performing the following steps, where _i_ is the accessor element index from zero (inclusive) to _count_ (exclusive): 1. Let _sparseCount_ be the number of replaced elements as defined by the `count` property of the `sparse` object. 2. Let _sparseIndices_ be the array of sparse indices and _sparseValues_ be the array of replacement elements as defined above; both arrays have _sparseCount_ elements. 3. If the _sparseIndices_ array contains the value of _i_, the value of the _i_-th accessor element is the _j_-th value of the _sparseValues_ array where _j_ is the index of the value of _i_ in the _sparseIndices_ array. + [NOTE] .Implementation Note ==== Binary search can be used for finding _i_ in the _sparseIndices_ array because the array is pre-sorted and does not contain duplicates. For the same reasons, the search can be avoided entirely if iterating over all accessor elements. ==== 4. Else if the _sparseIndices_ array does not contain the value of _i_, the value of the _i_-th accessor element is as if the `sparse` property is not defined. [[accessor-bounds]] ==== Accessor Bounds The `min` and `max` properties of the accessor are arrays that contain per-component minimum and maximum values, respectively. The length of these arrays **MUST** be equal to the number of accessor's components. If the accessor's component type is a non-normalized integer type, the minimum and maximum values stored as JSON decimal numbers **MUST** exactly match minimum and maximum values that would be accessed. If the accessor's component type is a normalized integer type, the minimum and maximum values stored as JSON decimal numbers **MUST** exactly match integer minimum and maximum values that would be accessed before they are converted to normalized floats. In other words, even if the accessor's `normalized` property is true, the bounds are stored as if it is false. If the accessor's component type is a floating-point type, the minimum and maximum values stored as JSON decimal numbers **SHOULD** exactly match minimum and maximum values that would be accessed. If they do not match exactly as-is, they **MUST** match exactly after being rounded to the accessor's specific floating-point component type. [NOTE] .Implementation Note ==== JSON usually implies double precision and some tools output JSON numbers that are not representable with single-precision floats but are close enough to match the accessed values after being rounded to the accessor's component type. Let's say the accessor's component type is _float32_ and the exact minimum value accessed with it is `1.2000000476837158203125`. The accessor's `min` property can be `1.2` in this case because rounding `1.2` to a single-precision float would produce the accessed value. It is therefore recommended to apply rounding to the `min` and `max` property values of floating-point accessors when loading glTF assets to avoid any potential runtime mismatches. If using ECMAScript, the `Math.fround` function could be used to round a JSON number to the nearest single-precision float. ==== [NOTE] .Implementation Note ==== The rules above imply that values of the `min` and `max` properties are always in the range of the accessor's component type and therefore implementations can internally use the same data type for the bounds as for the accessor elements. ==== If the accessor is sparse, its `min` and `max` properties correspond to the minimum and maximum component values after applying the sparse replacements. When neither `sparse` nor `bufferView` properties are defined, i.e., if all accessed elements are implicit zeros, the `min` and `max` properties **MAY** have any values representable with the accessor's component type. This is intended for use cases when data is supplied by external means (e.g., via extensions) and accessors are only used as format descriptors. Although the `min` and `max` properties are optional in general, accessors used for vertex positions or animation sampler inputs **MUST** define these properties. [[geometry]] == Geometry [[geometry-overview]] === Overview Any node **MAY** contain one mesh, defined in its `mesh` property. Mesh primitives of the mesh **MAY** have morph targets. The mesh **MAY** be skinned using information provided in the referenced `skin` object. [[meshes]] === Meshes [[meshes-overview]] ==== Overview A _mesh_ is defined as an array of _primitives_. A primitive object describes vertex attributes (data associated with each vertex), primitive topology (e.g., triangle list), vertex connectivity, and an optional material. Vertex attribute values are provided by accessors associated with vertex attribute semantics defined in the `attributes` property, primitive topology is defined in the enumerated `mode` property, and vertex connectivity is either directly derived from the order of attribute values, or it is defined by the index values provided by the accessor referenced by the `indices` property. The material used for rendering the primitive is defined by the `material` property. If the `material` property is not defined, then a <> is used. [NOTE] .Implementation Note ==== Splitting one mesh into several primitives can be useful to limit the amount of geometry rendered per a single draw call or to assign different materials to different parts of the mesh. ==== [NOTE] .Example ==== The following example defines a mesh containing one primitive with explicit vertex indices. [source,json] ---- { "meshes": [ { "primitives": [ { "attributes": { "NORMAL": 23, "POSITION": 22, "TANGENT": 24, "TEXCOORD_0": 25 }, "indices": 21, "material": 3, "mode": 4 } ] } ] } ---- ==== [[meshes-attributes]] ==== Vertex Attributes The `attributes` property is a JSON object that maps attribute semantic names to the indices of accessors providing the corresponding data. The `attributes` JSON object differs from the most of glTF JSON objects. In particular, it is treated as a flat key-value storage and thus it **MUST NOT** have any nested objects or values that are not accessor references. Properties of the `attributes` object **SHOULD** belong to one of the following categories as follows: - Attribute semantics defined in this Specification, i.e., `POSITION`, `NORMAL`, and `TANGENT`. - Indexed attribute semantics defined in this Specification, i.e., `TEXCOORD_n`, `COLOR_n`, `JOINTS_n`, and `WEIGHTS_n`, where `n` is an index placeholder. - Attribute semantics defined in glTF extensions using the `:` syntax, e.g., `EXT_my_extension:ATTRIBUTE`, provided that the corresponding extension is listed in the `extensionsUsed` property of the root glTF JSON object and that extension defines the attribute semantic used. - Attribute semantics that start with an underscore, e.g., `_TEMPERATURE`. These are considered application-specific. Attribute semantic names not belonging to any of these categories are out of scope of this Specification and **SHOULD NOT** be present in the `attributes` object. The following table provides attribute semantic names with their compatible accessor types and effective accessor component types for the attribute semantics defined in this Specification. Any combination of types and/or effective component types not present in this table **MUST NOT** be used for attribute semantics defined in this Specification. Future specification versions or extensions **MAY** define new attribute semantics as well as additional compatible types and/or effective component types for existing attribute semantics. Accessors of the _uint32_ effective component type **SHOULD NOT** be used for any vertex attribute semantic, including application-specific semantics, unless explicitly enabled by future specification versions or extensions. [options="header",cols="20%,20%,30%,30%"] |==== | Semantic Name | Accessor Type(s)| Effective Component Type(s) | Description | `POSITION` | VEC3 | _float32_ | Unitless XYZ vertex positions | `NORMAL` | VEC3 | _float32_ | Normalized XYZ vertex normals | `TANGENT` | VEC4 | _float32_ | XYZW vertex tangents where the XYZ portion is normalized, and the W component is a sign value (-1 or +1) indicating handedness of the tangent basis | `TEXCOORD_n` | VEC2 | _float32_ + _unorm8_ + _unorm16_ | ST texture coordinates | `COLOR_n` | VEC3 + VEC4 | _float32_ + _unorm8_ + _unorm16_ | RGB or RGBA vertex color | `JOINTS_n` | VEC4 | _uint8_ + _uint16_ | See <> | `WEIGHTS_n` | VEC4 | _float32_ + _unorm8_ + _unorm16_ | See <> |==== The `TEXCOORD_n`, `COLOR_n`, `JOINTS_n`, and `WEIGHTS_n` attribute semantic property names have the form `[semantic]_[index]`, e.g., `TEXCOORD_0`, `TEXCOORD_1`, `COLOR_0`. The following rules apply to the `[index]` syntax: - the `[index]` value **MUST** be a non-negative decimal integer; - the `[index]` value **MUST NOT** use leading zeroes for positive index values, e.g., `TEXCOORD_01` is invalid; - the `[index]` value **MUST NOT** have more than nine digits. [NOTE] ==== In other words, the `[index]` value is a decimal number from 0 to 999999999 without leading zeros. ==== Any attribute semantic name that starts with `TEXCOORD_`, `COLOR_`, `JOINTS_`, or `WEIGHTS_` characters that are not followed by a valid index value is invalid and **MUST NOT** be present. Index values used for the same attribute semantic within a single primitive **MUST NOT** have gaps, i.e., if an indexed attribute semantic name with index `n` is present for any positive `n`, then the same indexed attribute semantic name with index `n-1` **MUST** also be present. Skinning-related attribute semantics are always used together, so either both `JOINTS_n` and `WEIGHTS_n` names **MUST** be present or none for any given `n`. Among all possible indexed attribute semantic names, client implementations **SHOULD** support at least `TEXCOORD_0` and `TEXCOORD_1`, `COLOR_0`, and `JOINTS_0` with `WEIGHTS_0`. Attribute data accessed with floating-point accessors **MUST NOT** contain infinite or NaN values. Accessors for the `POSITION` semantic **MUST** have their bounds, i.e., `min` and `max` properties, defined. The W components provided by accessors for the `TANGENT` semantic **MUST** have the absolute value of stem:[1.0]. All components of each `COLOR_0` accessor element **MUST** be in the stem:[[0.0, 1.0\]] range. All attribute accessors for a given primitive **MUST** have the same number of elements, i.e., values of their `count` properties **MUST** be equal. [[meshes-topology]] ==== Topology and Connectivity Accessors referenced from the `attributes` object only provide attribute values without implying any specific topology or connectivity. Assembling the mesh primitive from those attribute values is controlled by the `indices` and `mode` properties. If the primitive's `indices` property is not defined, the mesh primitive is assembled by taking all elements of the attribute accessors in the order as they are provided by the accessors. The total number of the primitive's vertices is equal to the value of the `count` property of any attribute accessor referenced from the primitive's `attributes` object. [NOTE] .Implementation Note ==== In other words, the attribute values of the stem:[i]-th vertex of the mesh primitive are given by the stem:[i]-th elements of the corresponding accessors. For example, the position and normal of the first vertex of the mesh primitive are given by the first elements of the accessors referenced by the `POSITION` and `NORMAL` attributes, respectively; the position and normal of the second vertex of the mesh primitive are given by the second elements of those accessors, and so on. This kind of data layout is often called "`non-indexed`" because there is no indirection between the primitive's vertices and the attribute values. It is typically used for topology types that do not reuse the same vertex multiple times, such as point lists, line strips, or triangle strips (see below). ==== If the primitive's `indices` property is defined, it refers to the accessor that provides vertex indices, and the mesh primitive is assembled by taking the elements of the vertex indices accessor and fetching the corresponding elements of the attribute accessors in the order as they are provided by the vertex indices accessor. The total number of the primitive's vertices is equal to the value of the `count` property of the vertex indices accessor. [NOTE] .Implementation Note ==== In other words, the attribute values of the stem:[i]-th vertex of the mesh primitive are given by the stem:[k_i]-th elements of the corresponding accessors, where stem:[k_i] is the stem:[i]-th element value of the vertex indices accessor. Let's say the `POSITION` accessor provides five elements stem:[{ p_0, p_1, p_2, p_3, p_4 }], and the vertex indices accessor provides the following six elements stem:[{ 1, 2, 4, 2, 3, 4 }]. Then the mesh primitive has six vertices with positions stem:[{ p_1, p_2, p_4, p_2, p_3, p_4 }]. Note that the index values do not have to be unique or cover all possible accessor elements. If there are multiple vertex attributes, the same index values are used for all of them. For example, it is not possible to specify separate index values for positions and normals, so representing a cube made of twelve triangles with positions and flat normals would require 24 unique index values. This kind of data layout is often called "`indexed`" and it is typically used for topology types that reuse colocated vertices, such as triangle lists or line lists (see below). ==== The accessor referenced by the `indices` property **MUST** have the `"SCALAR"` type and any of the _uint8_, _uint16_, or _uint32_ effective component types. All values provided by the indices accessor **MUST** be less than the value of the `count` properties of the attribute accessors, i.e., each index value **MUST** refer to an existing vertex attribute value. Additionally, the accessor **MUST NOT** provide the maximum possible value for its component type, i.e., 255 for _uint8_, 65535 for _uint16_, or 4294967295 for _uint32_. [NOTE] .Implementation Note ==== The maximum values unconditionally trigger primitive restart in some graphics APIs and would require client implementations to rebuild the index data. ==== The rest of this section uses the "`number of the primitive's vertices`" term as defined above, i.e., the value of the `count` property of the vertex indices accessor if it is defined, or the value of the `count` property of the attribute accessors otherwise. In all cases, the number of the primitive's vertices **MUST** be positive. The primitive's vertices are connected based on the value of the enumerated `mode` property that defines the topology type of the mesh primitive. The following list provides all topology types defined by this specification. Most topology types have specific requirements for the number of the primitive's vertices. * **Point Lists** + Each vertex defines a single point primitive, according to the equation: {empty}:: stem:[P_i = { v_i }] * **Line Strips** + The number of the primitive's vertices **MUST** be greater than or equal to two. + One line primitive is defined by each vertex and the following vertex, according to the equation: {empty}:: stem:[P_i = { v_i, v_{i+1} }] * **Line Loops** + The number of the primitive's vertices **MUST** be greater than or equal to two. + Loops are the same as line strips except that a final segment is added from the final specified vertex to the first vertex. * **Line Lists** + The number of the primitive's vertices **MUST** be divisible by two. + Each consecutive pair of vertices defines a single line primitive, according to the equation: {empty}:: stem:[P_i = { v_{2i}, v_{2i+1} }] * **Triangle Lists** + The number of the primitive's vertices **MUST** be divisible by three. + Each consecutive set of three vertices defines a single triangle primitive, according to the equation: {empty}:: stem:[P_i = { v_{3i}, v_{3i+1}, v_{3i+2} }] * **Triangle Strips** + The number of the primitive's vertices **MUST** be greater than or equal to three. + One triangle primitive is defined by each vertex and the two vertices that follow it, according to the equation: {empty}:: stem:[P_i = { v_i, v_{i+(1+i%2)}, v_{i+(2-i%2)} }] * **Triangle Fans** + The number of the primitive's vertices **MUST** be greater than or equal to three. + Triangle primitives are defined around a shared common vertex, according to the equation: {empty}:: stem:[P_i = { v_{i+1}, v_{i+2}, v_0 }] Mesh geometry **SHOULD NOT** contain degenerate lines or triangles, i.e., lines or triangles that use the same vertex more than once per topology primitive. [[meshes-usage]] ==== Using Primitive Data When positions are not specified, client implementations **SHOULD** skip primitive's rendering unless its positions are provided by other means (e.g., by an extension). This applies to both indexed and non-indexed geometry. When normals are not specified, the mesh primitive has implicit flat normals. The provided tangents are ignored in this case and thus they **SHOULD NOT** be present. [NOTE] .Implementation Note ==== Flat normals mean that there is only one normal vector value per each rendered triangle. Depending on the mesh primitive connectivity, precomputing flat normals could require increasing the number of vertices and reassembling the mesh primitive. ==== When tangents are not specified and the material has a normal texture, the tangent vectors are calculated using default <> algorithms with the vertex positions, normals (specified or generated as mentioned above), and texture coordinates associated with the normal texture. Vertices of the same triangle **SHOULD** have equal W components of their `TANGENT` values. When vertices of the same triangle have different W values, its tangent space is considered undefined. The bitangent vectors are computed by taking the cross product of the XYZ components of the normal and tangent vectors and multiplying that cross product against the W component of the tangent. [latexmath] +++++ \vec{B} = (\vec{N_{xyz}} \times \vec{T_{xyz}}) \cdot T_w +++++ When a `COLOR_n` attribute uses an accessor of the `"VEC3"` type, its alpha component is assumed to have a value of stem:[1.0]. [[morph-targets]] ==== Morph Targets [[morph-targets-overview]] ===== Overview A mesh primitive **MAY** define _morph targets_. A morph target defines an alternative state of the primitive's vertex attribute values stored as differences (deltas) between the target vertex attribute values and the original vertex attributes values. Morph targets are specified in the primitive's `targets` array. The number of elements of the `targets` array defines the number of morph targets for the mesh primitive. Each element of that array is a non-empty JSON object that maps attribute semantic names to the indices of accessors providing the corresponding morph target difference values. Similarly to the primitive's `attributes` object, each element of the `targets` array is treated as a flat key-value storage and thus it **MUST NOT** have any nested objects or values that are not accessor references. Vertex attribute semantic names not present in the primitive's `attributes` object **MUST NOT** be present in its morph targets. Extension-defined and application-specific attribute semantics use the same naming conventions as for the primitive's `attributes` object. If a vertex attribute semantic name is not present in a morph target, the difference values for that vertex attribute semantic are zeros, i.e., the morph target does not modify values of that vertex attribute. [NOTE] .Implementation Note ==== This allows omitting zero-filled accessors and implies that different morph targets of the same primitive can modify different vertex attributes. ==== The number of morph targets of the mesh is the number of morph targets of the first mesh primitive that has morph targets. If a mesh has multiple mesh primitives, all mesh primitives of that mesh that define morph targets **MUST** have the same number of morph targets. [NOTE] .Implementation Note ==== When a mesh has a mix of mesh primitives with and without defined morph targets, the mesh primitives without morph targets can be interpreted as if they define the same number of morph targets with all difference values set to zeros. ==== Each morph target is applied to the mesh primitive according to the morph target's weight. When a mesh has multiple mesh primitives, the same morph target weights are used for all mesh primitives that have morph targets. The `weights` property of the mesh, if defined, provides an array of default morph target weights; the number of elements of the `weights` array **MUST** be equal to the number of the morph targets of the mesh. If the `weights` property is not defined, the default morph target weights are zeros. [NOTE] .Example ==== The following example shows a mesh with three primitives and two morph targets. The first primitive has no morph targets and thus its vertex attributes have only one state. The second primitive defines two morph targets with difference values for both vertex attribute semantics. The third primitive defines two morph targets and each morph target provides difference values for only one vertex attribute. The mesh also defines default morph target weights. [source,json] ---- { "primitives": [ { "attributes": { "POSITION": 0, "NORMAL": 1 } }, { "attributes": { "POSITION": 2, "NORMAL": 3 }, "targets": [ { "POSITION": 4, "NORMAL": 5 }, { "POSITION": 6, "NORMAL": 7 } ] }, { "attributes": { "POSITION": 8, "COLOR_0": 9 }, "targets": [ { "POSITION": 10 }, { "COLOR_0": 11 } ] } ], "weights": [ 0.5, 1.75 ] } ---- ==== [NOTE] .Implementation Note ==== A significant number of authoring and client implementations associate names with morph targets. While the glTF 2.0 specification currently does not provide a way to specify names, many tools add a `targetNames` array of strings to the `extras` object of the mesh containing the morph targets. The number of elements of the `targetNames` array is equal to the number of the morph targets of the mesh. ==== [[morph-targets-storage]] ===== Difference Data Storage The following table provides attribute semantic names that support morph targets with their compatible accessor types and effective accessor component types for the attribute semantics defined in this Specification. Any combination of types and/or effective component types not present in this table **MUST NOT** be used for attribute semantics defined in this Specification. Future specification versions or extensions **MAY** define new attribute semantics that support morph targets. Accessors of the _uint32_ effective component type **SHOULD NOT** be used for morph target differences, unless explicitly enabled by future specification versions or extensions. [options="header",cols="20%,20%,30%,30%"] |==== | Name | Accessor Type(s) | Effective Component Type(s) | Description | `POSITION` | VEC3 | _float32_ | XYZ vertex position displacements | `NORMAL` | VEC3 | _float32_ | XYZ vertex normal displacements | `TANGENT` | VEC3 | _float32_ | XYZ vertex tangent displacements | `TEXCOORD_n` | VEC2 | _float32_ + _snorm8_ + _snorm16_ + _unorm8_ + _unorm16_ | ST texture coordinate displacements | `COLOR_n` | VEC3 + VEC4 | _float32_ + _snorm8_ + _snorm16_ + _unorm8_ + _unorm16_ | RGB or RGBA color deltas |==== Client implementations **SHOULD** support at least three attributes -- `POSITION`, `NORMAL`, and `TANGENT` -- for morphing. Client implementations **MAY** optionally support morphed `TEXCOORD_n` and/or `COLOR_n` attributes. Morph data accessed with floating-point accessors **MUST NOT** contain infinite or NaN values. Accessors for the `POSITION` semantic **MUST** have their bounds, i.e., `min` and `max` properties, defined. Accessors for the `TANGENT` semantic do not have the W component for handedness since handedness cannot be displaced. All morph target accessors for a given primitive **MUST** have the same number of elements as the accessors providing original vertex attributes. [[morph-targets-usage]] ===== Applying Morph Data Let: - stem:[M] be the number of morph targets for a primitive and stem:[m] be the index of a morph target, stem:[m \in [0,M\)]; - stem:[A_i] be the original value of an attribute stem:[A] at a vertex index stem:[i]; - latexmath:[\Delta{A}_{m,i}] be the difference value of the attribute stem:[A] for the morph target stem:[m] at vertex index stem:[i]; - stem:[W_m] be the scalar weight for the stem:[m]-th morph target. Then the effective value of the vertex attribute stem:[A] for the vertex stem:[i] is computed as follows. [latexmath] +++++ A'_i = A_i + \sum_{m=0}^{M-1} \Delta{A}_{m,i} \cdot W_m +++++ If the accessor providing stem:[A_i] values has fewer components than are otherwise supported for the vertex attribute semantic, the omitted components are considered present with their implied values for the purpose of this computation. [NOTE] .Implementation Note ==== For example, original `COLOR_n` values stored with the `"VEC3"` type imply that their alpha values are stem:[1.0] and thus `COLOR_n` delta values provided by `"VEC4"` accessors are applied as if the original values also used the `"VEC4"` type. ==== If the accessor providing latexmath:[\Delta{A}_{m,i}] values has fewer components than are otherwise supported for the vertex attribute semantic, the omitted components are considered zeros for the purpose of this computation. [NOTE] .Implementation Note ==== For example, deltas for the `TANGENT` vertex attribute semantic do not contain the W component so morph targets do not modify tangent space handedness. Similarly, deltas for `COLOR_n` vertex attribute semantics provided via `"VEC3"` accessors do not modify alpha components. ==== These computations **SHOULD** use at least 32-bit floating-point precision. Displacements for `POSITION`, `NORMAL`, and `TANGENT` attributes are applied before any transformation matrices affecting the mesh vertices such as skinning or node transforms. If the mesh primitive does not specify normals, i.e., if they are implicitly flat, then each morph target also has flat normals, adjusted for the vertex position deltas if necessary. If the mesh primitive does specify normals but a morph target does not provide deltas for them, the normal vectors remain intact for that morph target. If the tangent vectors were calculated for the mesh primitive (because they were not provided or because the provided tangent vectors were ignored), the tangent vector deltas are also calculated for each morph target using the same algorithms and vertex positions, normals, and texture coordinates with the corresponding deltas applied. The provided tangent deltas are ignored in this case and thus they **SHOULD NOT** be present. If the mesh primitive does specify normals and tangents but a morph target does not provide tangent deltas, the tangent vectors remain intact for that morph target. Client implementations **SHOULD** clamp all `COLOR_0` attribute components to the to stem:[[0, 1\]] range after applying morph deltas. The number of morph targets is not limited. Client implementations **SHOULD** support at least eight morphed attributes. This means that they **SHOULD** support eight morph targets when each morph target has one attribute, four morph targets where each morph target has two attributes, or two morph targets where each morph target has three or four attributes. For assets that contain more morphed attributes than could be efficiently supported, client implementations **SHOULD** choose morph targets with the highest weights. [[skins]] === Skins [[skins-overview]] ==== Overview glTF 2.0 meshes support Linear Blend Skinning via skin objects, joint hierarchies, and designated vertex attributes. Skins are stored in the `skins` array of the asset. Each skin is defined by a **REQUIRED** `joints` property that lists the indices of nodes used as joints to pose the skin and an **OPTIONAL** `inverseBindMatrices` property that points to an accessor with inverse bind matrices data used to bring coordinates being skinned into the same space as each joint. The order of joints is defined by the `skin.joints` array and it **MUST** match the order of `inverseBindMatrices` accessor elements (when the latter is present). The `skeleton` property (if present) points to the node that is the common root of a joints hierarchy or to a direct or indirect parent node of the common root. [NOTE] .Implementation Note ==== Although the `skeleton` property is not needed for computing skinning transforms, it may be used to provide a specific "`pivot point`" for the skinned geometry. ==== An accessor referenced by `inverseBindMatrices` **MUST** have floating-point components of `"MAT4"` type. The number of elements of the accessor referenced by `inverseBindMatrices` **MUST** be greater than or equal to the number of `joints` elements. The fourth row of each matrix **MUST** be set to `[0.0, 0.0, 0.0, 1.0]`. The accessed matrices **MUST NOT** contain infinite or NaN values. [NOTE] .Implementation Note ==== The matrix defining how to pose the skin's geometry for use with the joints (also known as "`Bind Shape Matrix`") should be premultiplied to mesh data or to Inverse Bind Matrices. ==== [[joint-hierarchy]] ==== Joint Hierarchy The joint hierarchy used for controlling skinned mesh pose is simply the node hierarchy, with each node designated as a _joint_ by a reference from the `skin.joints` array. Each skin's joints **MUST** have a common parent node (direct or indirect) called _common root_, which may or may not be a joint node itself. When a skin is referenced by a node within a scene, the common root **MUST** belong to the same scene. [NOTE] .Implementation Note ==== A node object does not specify whether it is a joint. Client implementations may need to traverse the `skins` array first, marking each joint node. ==== A joint node **MAY** have other nodes attached to it, even a complete node sub graph with meshes. [NOTE] .Implementation Note ==== It's common to have an entire geometry attached to a joint node without having it being skinned (e.g., a sword attached to a hand). Note that the node transform is the local transform of the node relative to the joint, like any other node in the glTF node hierarchy as described in the <> section. ==== Only the joint transforms are applied to the skinned mesh; the transform of the skinned mesh node **MUST** be ignored. In the example below, the translation of `node_0` and the scale of `node_1` are applied while the translation of `node_3` and rotation of `node_4` are ignored. [source,json] ---- { "nodes": [ { "name": "node_0", "children": [ 1 ], "translation": [ 0.0, 1.0, 0.0 ] }, { "name": "node_1", "children": [ 2 ], "scale": [ 0.5, 0.5, 0.5 ] }, { "name": "node_2" }, { "name": "node_3", "children": [ 4 ], "translation": [ 1.0, 0.0, 0.0 ] }, { "name": "node_4", "mesh": 0, "rotation": [ 0.0, 1.0, 0.0, 0.0 ], "skin": 0 } ], "skins": [ { "inverseBindMatrices": 0, "joints": [ 1, 2 ], "skeleton": 1 } ] } ---- [[skinned-mesh-attributes]] ==== Skinned Mesh Attributes The skinned mesh **MUST** have vertex attributes that are used in skinning calculations. The `JOINTS_n` attribute data contains the indices of the joints from the corresponding `skin.joints` array that affect the vertex. The `WEIGHTS_n` attribute data defines the weights indicating how strongly the joint influences the vertex. To apply skinning, a transformation matrix is computed for each joint. Then, the per-vertex transformation matrices are computed as weighted linear sums of the joint transformation matrices. Note that per-joint inverse bind matrices (when present) **MUST** be applied before the base node transforms. In the following example, a mesh primitive defines `JOINTS_0` and `WEIGHTS_0` vertex attributes: [source,json] ---- { "meshes": [ { "name": "skinned-mesh_1", "primitives": [ { "attributes": { "JOINTS_0": 179, "NORMAL": 165, "POSITION": 163, "TEXCOORD_0": 167, "WEIGHTS_0": 176 }, "indices": 161, "material": 1, "mode": 4 } ] } ] } ---- The number of joints that influence one vertex is limited to 4 per set, so the referenced accessors **MUST** have `VEC4` type and following component types: * *`JOINTS_n`*: _uint8_ or _uint16_ * *`WEIGHTS_n`*: _float32_, or _unorm8_, or _unorm16_ The joint weights for each vertex **MUST NOT** be negative. Joints **MUST NOT** contain more than one non-zero weight for a given vertex. When the weights are stored using _float_ component type, their linear sum **SHOULD** be as close as reasonably possible to `1.0` for a given vertex. When the weights are stored using _unorm8_, or _unorm16_ component types, their linear sum before normalization **MUST** be `255` or `65535`, respectively. Without these requirements, vertices would be deformed significantly because the weight error would get multiplied by the joint position. For example, an error of `1/255` in the weight sum would result in an unacceptably large difference in the joint position. [NOTE] .Implementation Note ==== The threshold in the official validation tool is set to `2e-7` times the number of non-zero weights per vertex. ==== [NOTE] .Implementation Note ==== Since the allowed threshold is much lower than minimum possible step for quantized component types, weight sum should be renormalized after quantization. ==== When any of the vertices are influenced by more than four joints, the additional joint and weight information are stored in subsequent sets. For example, `JOINTS_1` and `WEIGHTS_1` if present will reference the accessor for up to 4 additional joints that influence the vertices. For a given primitive, the number of `JOINTS_n` attribute sets **MUST** be equal to the number of `WEIGHTS_n` attribute sets. Client implementations **MAY** support only a single set of up to four weights and joints, however not supporting all weight and joint sets present in the file may have an impact on the asset's animation. All joint values **MUST** be within the range of joints in the skin. Unused joint values (i.e., joints with a weight of zero) **SHOULD** be set to zero. [[instantiation]] === Instantiation A mesh is instantiated by a node's `mesh` property. The same mesh **MAY** be instantiated by multiple nodes, which **MAY** specify different transforms, morph weights, and/or skins. [NOTE] .Implementation Note ==== This example instantiates the same mesh twice: without any transforms and with the specified translation. [source,json] ---- { "nodes": [ { "mesh": 11 }, { "mesh": 11, "translation": [ -20, -1, 0 ] } ] } ---- ==== After applying the node's global transform, mesh vertex position values are meters. When a mesh primitive uses any triangle-based topology (i.e., _triangle lists_, _triangle strips_, or _triangle fans_), the determinant of the node's global transform defines the winding order of the instantiated mesh primitives: - if the determinant is a positive value, the front-facing triangles use the counterclockwise order, and the back-facing triangles use the clockwise order; - if the determinant is a negative value, the front-facing triangles use the clockwise order, and the back-facing triangles use the counterclockwise order. [NOTE] .Implementation Note ==== Switching the winding order to clockwise enables mirroring geometry via negative scale transforms. ==== When an instantiated mesh has morph targets, the node's `weights` array property overrides mesh's default morph target weights; the number of elements of that array **MUST** be equal to the number of the morph targets of the mesh. If the `weights` property is not defined, the mesh is instantiated with its default morph target weights. [NOTE] .Implementation Note ==== Default morph target weights of a mesh are either specified by the `weights` property of the mesh, if that property is defined, or zeros otherwise. ==== The example below instantiates a Morph Target with non-default weights. [source,json] ---- { "nodes": [ { "mesh": 11, "weights": [0, 0.5] } ] } ---- A skin is instantiated within a node using a combination of the node's `mesh` and `skin` properties. The mesh for a skin instance is defined in the `mesh` property. The `skin` property contains the index of the skin to instance. The following example shows a skinned mesh instance: a skin object, a node with a skinned mesh, and two joint nodes. [source,json] ---- { "skins": [ { "inverseBindMatrices": 29, "joints": [1, 2] } ], "nodes": [ { "name":"Skinned mesh node", "mesh": 0, "skin": 0 }, { "name":"Skeleton root joint", "children": [2], "rotation": [ 0, 0, 0.7071067811865475, 0.7071067811865476 ], "translation": [ 4.61599, -2.032e-06, -5.08e-08 ] }, { "name":"Head", "translation": [ 8.76635, 0, 0 ] } ] } ---- [[texture-data]] == Texture Data [[texture-data-overview]] === Overview glTF 2.0 separates texture access into three distinct types of objects: Textures, Images, and Samplers. [[textures]] === Textures glTF 2.0 supports only static 2D textures. Textures are stored in the asset's `textures` array. A texture is defined by the `source` and `sampler` properties that define the texture's image and sampler, respectively. The image provides texture's data, and the sampler defines wrapping and filtering modes for the texture. [NOTE] .Example ==== [source,json] ---- { "textures": [ { "sampler": 0, "source": 2 } ] } ---- ==== When the `source` property of a texture is undefined, the texture's image **SHOULD** be provided by an extension or application-specific means, otherwise the texture's image is undefined and using the texture could result in undefined rendering results. [NOTE] .Implementation Note ==== Client implementations could render such textures with a predefined placeholder image or being filled with an error color (e.g., magenta). ==== When the `sampler` property of a texture is undefined, a sampler with repeat wrapping in both directions and implementation-dependent default texture filtering is used. [NOTE] .Implementation Note ==== This behavior matches the default state of a sampler object with all its properties omitted. ==== The origin of the texture coordinates (0, 0) corresponds to the upper left corner of a texture image. This is illustrated in the following figure, where the respective coordinates are shown for all four corners of a normalized texture space: .Normalized Texture Coordinates image::figures/texcoords.svg[pdfwidth=3in,align=left] [[images]] === Images [[images-overview]] ==== Overview Image objects are stored in the asset's `images` array. An image object represents a reference to an image resource. The glTF 2.0 specification supports only image resources stored in PNG and JPEG image formats (further defined below). Support for other image formats **MAY** be added by extensions or future Specification versions. The image reference is given either by the `bufferView` or `uri` properties; exactly one of these two properties **MUST** be defined. If the `uri` property is defined, it either provides a URI (or IRI) to the external image file, or embeds the image in the URI using the Data URI scheme. When the Data URI scheme is used for image storage, the `mediatype` field of the URI **MUST** be set to the media type of the embedded image data. If the URI cannot be resolved, the image's content is undefined and handling of such image objects is implementation-dependent. If the `bufferView` property is defined, the image data is sourced from the referenced buffer view. The buffer view **MUST NOT** have its `byteStride` or `target` properties defined. [NOTE] .Implementation Note ==== The entire buffer view is used as the image data source, i.e., there is no way to specify start or end byte offsets. ==== If the `mimeType` property is defined, it explicitly specifies the media type of the image object and the format of the image data **MUST** match the value of that property. The `mimeType` property **MUST** be defined if the image is sourced from a buffer view. [NOTE] .Example ==== This example shows an image pointing to an external PNG image file and another image referencing a buffer view with JPEG data. [source,json] ---- { "images": [ { "uri": "duck.png" }, { "bufferView": 14, "mimeType": "image/jpeg" } ] } ---- ==== Client implementations **MAY** need to manually determine image media types. In such a case, the following table **SHOULD** be used to derive the media type based in the values of the first few bytes of the resource. [options="header"] |==== | Media Type | Pattern Length | Pattern Bytes | `image/png` | 8 | `0x89 0x50 0x4E 0x47 0x0D 0x0A 0x1A 0x0A` | `image/jpeg` | 3 | `0xFF 0xD8 0xFF` |==== Any colorspace information (such as ICC profiles, intents, gamma values, etc.) found in images **MUST** be ignored. The effective transfer function (encoding) is defined by a glTF object that refers to the image (in most cases it's a texture that is used by a material). The effective color gamut for all images that are used as color textures is ITU-R BT.709. [NOTE] .Web Implementation Note ==== To ignore embedded colorspace information when using WebGL API, set `UNPACK_COLORSPACE_CONVERSION_WEBGL` flag to `NONE`. To ignore embedded colorspace information when using ImageBitmap API, set `colorSpaceConversion` option to `none`. ==== Regardless of the storage format, images always have four color channels: red, green, blue, and alpha. Format-specific sections below define how these four channels correspond to the stored image channels. [[images-png]] ==== PNG Support <> images **MUST** use the <> Media Type. PNG images **SHOULD** use the `.png` file extension when stored as separate files. If a PNG image does not have an explicit alpha channel but includes a `tRNS` chunk, the image's alpha channel is controlled by the content of that chunk as defined in the PNG specification. If a PNG image has neither an explicit alpha channel, nor a `tRNS` chunk, the image is fully opaque, i.e., its alpha channel has a uniform value of stem:[1.0]. If a PNG image uses a "`Greyscale`" or "`Greyscale with alpha`" color types, the values of the red, green, and blue channels of the image are equal to the "`Grey`" samples. The following PNG features **SHOULD NOT** be present and **SHOULD** be ignored by client implementations if encountered: - colorspace and/or display information (`cHRM`, `gAMA`, `iCCP`, `sBIT`, `sRGB`, `cICP`, `mDCV`, and/or `cLLI` chunks); - non-square pixel aspect ratio (`pHYs` chunk); - Exchangeable Image File Profile (`eXIf` chunk or embedded in `tEXt`, `iTXt`, or `zTXt` chunks); - animations (`acTL`, `fcTL`, `fdAT` chunks). [NOTE] .Rationale ==== These features are not universally implemented, can conflict with glTF definitions, and thus they may severely impact an asset's portability, especially, Exif's "`Orientation`". ==== [NOTE] .Implementation Note ==== Chunk presence does not always mean that the corresponding feature is used, see the PNG specification for more information on how PNG chunks are interpreted. ==== [[images-jpeg]] ==== JPEG Support <> images **MUST** use the <> Media Type. JPEG images **SHOULD** use the `.jpg` or `.jpeg` file extensions when stored as separate files. JPEG images **MUST** conform to the <>, namely: - JPEG images **MUST** have an APP~0~ marker containing the null-terminated string `JFIF` immediately after the SOI marker; - JPEG images **MUST** have one or three image components with 8 bits per component; - JPEG images that have three image components **MUST** use the _YC~B~C~R~_ color model and ordered such that _Y_ is the first component, _C~B~_ is the second component, and _C~R~_ is the third component. + [NOTE] .Implementation Note ==== The _YC~B~C~R~_ color model is implicitly converted to RGB as defined in JFIF when the image is decoded. ==== If a JPEG image has only one image component (called _Y_ in JFIF), the values of the red, green, and blue channels of the image are equal to the _Y_ component. A JPEG image is always fully opaque, its alpha channel has a uniform value of stem:[1.0]. Additionally, for glTF compatibility the following restrictions apply. - JPEG images **MUST NOT** contain any start-of-frame markers other than SOF~0~, SOF~1~, or SOF~2~. + [NOTE] .Rationale ==== This restriction implies that JPEG images cannot use arithmetic coding, lossless process, or hierarchical mode of operation. These features are not widely supported and images that rely on them would not be portable. ==== - JPEG images **MUST NOT** contain DNL markers. + [NOTE] .Rationale ==== This restriction ensures that the image's height is equal to the number of lines set in the start-of-frame marker. DNL markers are not widely supported and images that rely on them would not be portable. ==== - JPEG images **SHOULD NOT** contain embedded ICC profiles. If present, they **MUST** be ignored by client implementations. - JPEG images **SHOULD NOT** contain <> data. If present, it **SHOULD** be ignored by client implementations. + [NOTE] .Rationale ==== JFIF and Exif formats are mutually exclusive. Some of Exif features, e.g., "`Orientation`", may severely impact an asset's portability, . ==== [[samplers]] === Samplers ==== Overview Samplers are stored in the `samplers` array of the asset. Each sampler specifies filtering and wrapping modes. The sampler properties use integer enums defined in the <>. Client implementations **SHOULD** follow specified filtering modes. When the latter are undefined, client implementations **MAY** set their own default texture filtering settings. Client implementations **MUST** follow specified wrapping modes. ==== Filtering Filtering modes control texture's magnification and minification. Magnification modes include: * _Nearest_. For each requested texel coordinate, the sampler selects a texel with the nearest coordinates. This process is sometimes called "`nearest neighbor`". * _Linear_. For each requested texel coordinate, the sampler computes a weighted sum of several adjacent texels. This process is sometimes called "`bilinear interpolation`". Minification modes include: * _Nearest_. For each requested texel coordinate, the sampler selects a texel with the nearest (in Manhattan distance) coordinates from the original image. This process is sometimes called "`nearest neighbor`". * _Linear_. For each requested texel coordinate, the sampler computes a weighted sum of several adjacent texels from the original image. This process is sometimes called "`bilinear interpolation`". * _Nearest-mipmap-nearest_. For each requested texel coordinate, the sampler first selects one of pre-minified versions of the original image, and then selects a texel with the nearest (in Manhattan distance) coordinates from it. * _Linear-mipmap-nearest_. For each requested texel coordinate, the sampler first selects one of pre-minified versions of the original image, and then computes a weighted sum of several adjacent texels from it. * _Nearest-mipmap-linear_. For each requested texel coordinate, the sampler first selects two pre-minified versions of the original image, selects a texel with the nearest (in Manhattan distance) coordinates from each of them, and performs final linear interpolation between these two intermediate results. * _Linear-mipmap-linear_. For each requested texel coordinate, the sampler first selects two pre-minified versions of the original image, computes a weighted sum of several adjacent texels from each of them, and performs final linear interpolation between these two intermediate results. This process is sometimes called "`trilinear interpolation`". To properly support mipmap modes, client implementations **SHOULD** generate mipmaps at runtime. When runtime mipmap generation is not possible, client implementations **SHOULD** override the minification filtering mode as follows: [options="header"] |==== | Mipmap minification mode | Fallback mode | _Nearest-mipmap-nearest_ + _Nearest-mipmap-linear_ | _Nearest_ | _Linear-mipmap-nearest_ + _Linear-mipmap-linear_ | _Linear_ |==== ==== Wrapping Per-vertex texture coordinates, which are provided via `TEXCOORD_n` attribute values, are normalized for the image size (not to confuse with the `normalized` accessor property, the latter refers only to data encoding). That is, the texture coordinate value of `(0.0, 0.0)` points to the beginning of the first (upper-left) image pixel, while the texture coordinate value of `(1.0, 1.0)` points to the end of the last (lower-right) image pixel. Sampler's wrapping modes define how to handle texture coordinates that are negative or greater than or equal to `1.0`, independently for both directions. Supported modes include: * _Repeat_. Only the fractional part of texture coordinates is used. + [NOTE] .Example ==== `2.2` maps to `0.2`; `-0.4` maps to `0.6`. ==== * _Mirrored Repeat_. This mode works as _repeat_ but flips the direction when the integer part (truncated towards −∞) is odd. + [NOTE] .Example ==== `2.2` maps to `0.2`; `-0.4` is treated as `0.4`. ==== * _Clamp to edge_. Texture coordinates with values outside the image are clamped to the closest existing image texel at the edge. ==== Example The following example defines a sampler with _linear_ magnification filtering, _linear-mipmap-linear_ minification filtering, and _repeat_ wrapping in both directions. [source,json] ---- { "samplers": [ { "magFilter": 9729, "minFilter": 9987, "wrapS": 10497, "wrapT": 10497 } ] } ---- ==== Non-power-of-two Textures Client implementations **SHOULD** resize non-power-of-two textures (so that their horizontal and vertical sizes are powers of two) when running on platforms that have limited support for such texture dimensions. [NOTE] .Implementation Note ==== Specifically, if the `sampler` the texture references: * has a wrapping mode (either `wrapS` or `wrapT`) equal to _repeat_ or _mirrored repeat_, or * has a minification filter (`minFilter`) that uses mipmapping. ==== [[materials]] == Materials [[materials-overview]] === Overview glTF defines materials using a common set of parameters that are based on widely used material representations from Physically Based Rendering (PBR). Specifically, glTF uses the metallic-roughness material model. Using this declarative representation of materials enables a glTF file to be rendered consistently across platforms. .Physically Based Rendering Example image::figures/materials.svg[pdfwidth=6in,align=left] [[metallic-roughness-material]] === Metallic-Roughness Material All parameters related to the metallic-roughness material model are defined under the `pbrMetallicRoughness` property of `material` object. The following example shows how to define a gold-like material using the metallic-roughness parameters: [source,json] ---- { "materials": [ { "name": "gold", "pbrMetallicRoughness": { "baseColorFactor": [ 1.000, 0.766, 0.336, 1.0 ], "metallicFactor": 1.0, "roughnessFactor": 0.0 } } ] } ---- The metallic-roughness material model is defined by the following properties: * _base color_ - The base color of the material. * _metalness_ - The metalness of the material; values range from `0.0` (non-metal) to `1.0` (metal); see <> for the interpretation of intermediate values. * _roughness_ - The roughness of the material; values range from `0.0` (smooth) to `1.0` (rough). The _base color_ has two different interpretations depending on the value of _metalness_. When the material is a metal, the base color is the specific measured reflectance value at normal incidence (F0). For a non-metal the base color represents the reflected diffuse color of the material. In this model it is not possible to specify a F0 value for non-metals, and a linear value of 4% (0.04) is used. The value for each property **MAY** be defined using factors and/or textures (e.g., `baseColorTexture` and `baseColorFactor`). If a texture is not given, all respective texture components within this material model have a value of `1.0`. The factor value is a linear multiplier for the corresponding texture components. A texture binding is defined by an `index` of a _texture_ object and an optional index of texture coordinates. The following example shows a material that uses a texture for its _base color_ property. [source,json] ---- { "materials": [ { "pbrMetallicRoughness": { "baseColorTexture": { "index": 0, "texCoord": 1 }, } } ], "textures": [ { "source": 0 } ], "images": [ { "uri": "base_color.png" } ] } ---- The _base color_ texture **MUST** contain 8-bit values encoded with the <> so RGB values **MUST** be decoded to real linear values before they are used for any computations. To achieve correct filtering, the transfer function **SHOULD** be decoded before performing linear interpolation. The textures for _metalness_ and _roughness_ properties are packed together in a single texture called `metallicRoughnessTexture`. Its _green_ channel contains roughness values and its _blue_ channel contains metalness values. This texture **MUST** be encoded with linear transfer function and **MAY** use more than 8 bits per channel. For example, assume an 8-bit RGBA value of `[64, 124, 231, 255]` is sampled from `baseColorTexture` and assume that `baseColorFactor` is given as `[0.2, 1.0, 0.7, 1.0]`. Then, the final _base color_ value would be (after decoding the transfer function and multiplying by the factor) [source,c] ---- [0.051 * 0.2, 0.202 * 1.0, 0.799 * 0.7, 1.0 * 1.0] = [0.0102, 0.202, 0.5593, 1.0] ---- In addition to the material properties, if a primitive specifies a vertex color using the attribute semantic property `COLOR_0`, then this value acts as an additional linear multiplier to _base color_. Implementations of the bidirectional reflectance distribution function (BRDF) itself **MAY** vary based on device performance and resource constraints. See <> for more details on the BRDF calculations. [[additional-textures]] === Additional Textures The material definition also provides for additional textures that **MAY** also be used with the metallic-roughness material model as well as other material models, which could be provided via glTF extensions. The following additional textures are supported: - *normal* : A tangent space normal texture. The texture encodes XYZ components of a normal vector in tangent space as RGB values stored with linear transfer function. Normal textures **SHOULD NOT** contain _alpha_ channel as it not used anyway. After dequantization, texel values **MUST** be mapped as follows: _red_ [0.0 .. 1.0] to X [-1 .. 1], _green_ [0.0 .. 1.0] to Y [-1 .. 1], _blue_ (0.5 .. 1.0] maps to Z (0 .. 1]. Normal textures **SHOULD NOT** contain blue values less than or equal to `0.5`. + [NOTE] .Implementation Note ==== This mapping is usually implemented as `sampledValue * 2.0 - 1.0`. ==== + The texture binding for normal textures **MAY** additionally contain a scalar `scale` value that linearly scales X and Y components of the normal vector. + Normal vectors **MUST** be normalized before being used in lighting equations. When scaling is used, vector normalization happens after scaling. - *occlusion* : The occlusion texture; it indicates areas that receive less indirect lighting from ambient sources. Direct lighting is not affected. The _red_ channel of the texture encodes the occlusion value, where `0.0` means fully-occluded area (no indirect lighting) and `1.0` means not occluded area (full indirect lighting). Other texture channels (if present) do not affect occlusion. + The texture binding for occlusion maps **MAY** optionally contain a scalar `strength` value that is used to reduce the occlusion effect. When present, it affects the occlusion value as `1.0 + strength * (occlusionTexture - 1.0)`. - *emissive* : The emissive texture and factor control the color and intensity of the light being emitted by the material. The texture **MUST** contain 8-bit values encoded with the <> so RGB values **MUST** be decoded to real linear values before they are used for any computations. To achieve correct filtering, the transfer function **SHOULD** be decoded before performing linear interpolation. + For implementations where a physical light unit is needed, the units for the multiplicative product of the emissive texture and factor are candela per square meter (**cd / m^2^**), sometimes called _nits_. + [NOTE] .Implementation Note ==== Because the value is specified per square meter, it indicates the brightness of any given point along the surface. However, the exact conversion from physical light units to the brightness of rendered pixels requires knowledge of the camera's exposure settings, which are left as an implementation detail unless otherwise defined by a glTF extension. Many rendering engines simplify this calculation by assuming that an emissive factor of `1.0` results in a fully exposed pixel. ==== The following example shows a material that is defined using `pbrMetallicRoughness` parameters as well as additional textures: [source,json] ---- { "materials": [ { "name": "Material0", "pbrMetallicRoughness": { "baseColorFactor": [ 0.5, 0.5, 0.5, 1.0 ], "baseColorTexture": { "index": 1, "texCoord": 1 }, "metallicFactor": 1, "roughnessFactor": 1, "metallicRoughnessTexture": { "index": 2, "texCoord": 1 } }, "normalTexture": { "scale": 2, "index": 3, "texCoord": 1 }, "emissiveFactor": [ 0.2, 0.1, 0.0 ] } ] } ---- If a client implementation is resource-bound and cannot support all the textures defined it **SHOULD** support these additional textures in the following priority order. Resource-bound implementations **SHOULD** drop textures from the bottom to the top. [options="header",cols="20%,80%"] |==== | Texture | Rendering impact when feature is not supported | Normal | Geometry will appear less detailed than authored. | Occlusion | Model will appear brighter in areas that are intended to be darker. | Emissive | Model with lights will not be lit. For example, the headlights of a car model will be off instead of on. |==== [[alpha-coverage]] === Alpha Coverage The `alphaMode` property defines how the alpha value is interpreted. The alpha value is taken from the fourth component of the _base color_ for metallic-roughness material model. `alphaMode` can be one of the following values: * `OPAQUE` - The rendered output is fully opaque and any alpha value is ignored. * `MASK` - The rendered output is either fully opaque or fully transparent depending on the alpha value and the specified _alpha cutoff_ value; the exact appearance of the edges **MAY** be subject to implementation-specific techniques such as "`Alpha-to-Coverage`". + [NOTE] .Note ==== This mode is used to simulate geometry such as tree leaves or wire fences. ==== * `BLEND` - The rendered output is combined with the background using the "`over`" operator as described in <>. + [NOTE] .Note ==== This mode is used to simulate geometry such as gauze cloth or animal fur. ==== When `alphaMode` is set to `MASK` the `alphaCutoff` property specifies the cutoff threshold. If the alpha value is greater than or equal to the `alphaCutoff` value then it is rendered as fully opaque, otherwise, it is rendered as fully transparent. `alphaCutoff` value is ignored for other modes. [NOTE] .Implementation Note for Real-Time Rasterizers ==== Real-time rasterizers typically use depth buffers and mesh sorting to support alpha modes. The following describe the expected behavior for these types of renderers. * `OPAQUE` - A depth value is written for every pixel and mesh sorting is not required for correct output. * `MASK` - A depth value is not written for a pixel that is discarded after the alpha test. A depth value is written for all other pixels. Mesh sorting is not required for correct output. * `BLEND` - Support for this mode varies. There is no perfect and fast solution that works for all cases. Client implementations should try to achieve the correct blending output for as many situations as possible. Whether depth value is written or whether to sort is up to the implementation. For example, implementations may discard pixels that have zero or close to zero alpha value to avoid sorting issues. ==== [[double-sided]] === Double Sided The `doubleSided` property specifies whether the material is double sided. When this value is false, back-face culling is enabled, i.e., only front-facing triangles are rendered. When this value is true, back-face culling is disabled, triangles are rendered from both sides, and the normal vectors of the back-facing triangles are negated. [[default-material]] === Default Material The default material, used when a mesh does not specify a material, is defined to be a material with no properties specified. All the default values of <> apply. [NOTE] .Implementation Note ==== This material does not emit light and will be black unless some lighting is present in the scene. ==== [[point-and-line-materials]] === Point and Line Materials This specification does not define size or style of non-triangular primitives (such as _points_ or _lines_), and applications **MAY** use various techniques to render these primitives as appropriate. However, the following conventions are **RECOMMENDED** for consistency: * Points and Lines **SHOULD** have widths of 1px in viewport space. * Points or Lines with `NORMAL` and `TANGENT` attributes **SHOULD** be rendered with standard lighting including normal textures. * Points or Lines with `NORMAL` but without `TANGENT` attributes **SHOULD** be rendered with standard lighting but ignoring any normal textures on the material. * Points or Lines with no `NORMAL` attribute **SHOULD** be rendered without lighting and instead use the sum of the _base color_ value (as defined above, multiplied by `COLOR_0` when present) and the _emissive_ value. [[cameras]] == Cameras [[cameras-overview]] === Overview Cameras are stored in the asset's `cameras` array. Each camera defines a `type` property that designates the type of projection (perspective or orthographic), and either a `perspective` or `orthographic` property that defines the details. A camera is instantiated within a node using the `camera` property of the node. A camera object defines the projection matrix that transforms scene coordinates from the view space to the clip space. A node containing the `camera` property instantiates a camera with that projection matrix and defines the associated view matrix that transforms scene coordinates from the global space to the view space. [[view-matrix]] === View Matrix A camera is defined such that the local +X axis is to the right, the "`lens`" looks towards the local -Z axis, and the top of the camera is aligned with the local +Y axis. The node that instantiates a camera defines the view matrix iff its global transformation matrix conforms to all of the following conditions: - none of the first three columns is all-zeros; - the determinant of the upper-left 3x3 matrix formed by dividing each matrix element by the length of its column is close to positive one. [NOTE] .Implementation Note ==== These conditions ensure that the rotation component can be unambiguously extracted from the matrix assuming that all scale factors are positive. ==== The view matrix is then derived from the rotation and translation components of the node's transformation matrix. If the node's global transform is identity, the location of the camera is at the origin. If the node's global transformation matrix does not conform to any of the conditions above, the view matrix is undefined. [[projection-matrices]] === Projection Matrices [[projection-matrices-overview]] ==== Overview The projection can be perspective or orthographic. There are two subtypes of perspective projections: finite and infinite. When the `zfar` property is undefined, the camera defines an infinite projection. Otherwise, the camera defines a finite projection. [NOTE] .Example ==== The following example defines two perspective cameras with supplied values for Y field of view, aspect ratio, and clipping information. [source,json] ---- { "cameras": [ { "name": "Finite perspective camera", "type": "perspective", "perspective": { "aspectRatio": 1.5, "yfov": 0.660593, "zfar": 100, "znear": 0.01 } }, { "name": "Infinite perspective camera", "type": "perspective", "perspective": { "aspectRatio": 1.5, "yfov": 0.660593, "znear": 0.01 } } ] } ---- ==== The effective projection matrices would heavily depend on the implementation details, in particular the clip space conventions, such as the range (zero-to-one or minus-one-to-one) and direction of the Z axis. [NOTE] .Implementation Note ==== When the aspect ratio derived from the camera properties (regardless of the camera type) does not match the aspect ratio of the current viewport, client implementations are free to use any context-appropriate content adaptation, such as cropping and/or uniform scaling. It is generally not recommended to perform non-uniform scaling ("`stretching`") as that would distort the content. ==== [[infinite-perspective-projection]] ==== Infinite Perspective Projection The infinite perspective projection is defined by the effective aspect ratio (width over height) and the `yfov` and `znear` properties that represent the vertical field of view in radians and the distance to the near clipping plane, respectively. The far clipping plane is at infinity. If the `aspectRatio` property is present, it defines the effective aspect ratio. If the `aspectRatio` property is not present, the effective aspect ratio is equal to the aspect ratio of the viewport. The values of the properties mentioned in this section **MUST** be greater than zero. [[finite-perspective-projection]] ==== Finite Perspective Projection The finite perspective projection is defined by the effective aspect ratio (width over height) and the `yfov`, `znear`, and `zfar` properties that represent the vertical field of view in radians, the distance to the near clipping plane, and the distance to the far clipping plane, respectively. If the `aspectRatio` property is present, it defines the effective aspect ratio. If the `aspectRatio` property is not present, the effective aspect ratio is equal to the aspect ratio of the viewport. The values of the properties mentioned in this section **MUST** be greater than zero. The value of the `zfar` property **MUST** be greater than the value of the `znear` property. [[orthographic-projection]] ==== Orthographic Projection The orthographic projection is defined by the `xmag`, `ymag`, `znear`, and `zfar` properties that represent the horizontal magnification factor (half of the orthographic width), the vertical magnification factor (half of the orthographic height), the distance to the near clipping plane, and the distance to the far clipping plane, respectively. The values of the `xmag` and `ymag` properties **SHOULD** be positive and they **MUST NOT** be zeros. The value of the `znear` property **MUST** greater than or equal to zero. The value of the `zfar` property **MUST** be greater than zero and greater than the value of the `znear` property. [[animations]] == Animations glTF supports articulated and skinned animation via key frame animations of nodes' transforms. Key frame data is stored in buffers and referenced in animations using accessors. glTF 2.0 also supports animation of instantiated morph targets in a similar fashion. [NOTE] .Note ==== glTF 2.0 only supports animating node transforms and morph target weights. Extensions or a future version of the specification may support animating arbitrary properties, such as material colors and texture transformation matrices. ==== [NOTE] .Note ==== glTF 2.0 defines only storage of animation keyframes, so this specification doesn't define any runtime behavior, such as: order of playing, auto-start, loops, mapping of timelines, etc. When loading a glTF 2.0 asset, client implementations may select an animation entry and pause it on the first frame, play it automatically, or ignore all animations until further user requests. When a playing animation is stopped, client implementations may reset the scene to the initial state or freeze it at the current frame. ==== [NOTE] .Implementation Note ==== glTF 2.0 does not specifically define how an animation will be used when imported but, as a best practice, it is recommended that each animation is self-contained as an action. For example, "`Walk`" and "`Run`" animations might each contain multiple channels targeting a model's various bones. The client implementation may choose when to play any of the available animations. ==== All animations are stored in the `animations` array of the asset. An animation is defined as a set of channels (the `channels` property) and a set of samplers that specify accessors with key frame data and interpolation method (the `samplers` property). The following examples show the expected usage of animations. [source,json] ---- { "animations": [ { "name": "Animate all properties of one node with different samplers", "channels": [ { "sampler": 0, "target": { "node": 1, "path": "rotation" } }, { "sampler": 1, "target": { "node": 1, "path": "scale" } }, { "sampler": 2, "target": { "node": 1, "path": "translation" } } ], "samplers": [ { "input": 4, "interpolation": "LINEAR", "output": 5 }, { "input": 4, "interpolation": "LINEAR", "output": 6 }, { "input": 4, "interpolation": "LINEAR", "output": 7 } ] }, { "name": "Animate two nodes with different samplers", "channels": [ { "sampler": 0, "target": { "node": 0, "path": "rotation" } }, { "sampler": 1, "target": { "node": 1, "path": "rotation" } } ], "samplers": [ { "input": 0, "interpolation": "LINEAR", "output": 1 }, { "input": 2, "interpolation": "LINEAR", "output": 3 } ] }, { "name": "Animate two nodes with the same sampler", "channels": [ { "sampler": 0, "target": { "node": 0, "path": "rotation" } }, { "sampler": 0, "target": { "node": 1, "path": "rotation" } } ], "samplers": [ { "input": 0, "interpolation": "LINEAR", "output": 1 } ] }, { "name": "Animate a node rotation channel and the weights of a Morph Target it instantiates", "channels": [ { "sampler": 0, "target": { "node": 1, "path": "rotation" } }, { "sampler": 1, "target": { "node": 1, "path": "weights" } } ], "samplers": [ { "input": 4, "interpolation": "LINEAR", "output": 5 }, { "input": 4, "interpolation": "LINEAR", "output": 6 } ] } ] } ---- _Channels_ connect the output values of the key frame animation to a specific node in the hierarchy. A channel's `sampler` property contains the index of one of the samplers present in the containing animation's `samplers` array. The `target` property is an object that identifies which node to animate using its `node` property, and which property of the node to animate using `path`. Non-animated properties **MUST** keep their values during animation. When the `node` property is not defined, the animation channel is ignored unless target information is provided by other means, such as extensions. When the `node` property is defined, valid values for the `path` property include `"translation"`, `"rotation"`, `"scale"`, and `"weights"`. If the node has the `matrix` property defined, its TRS properties **MUST NOT** be targeted, i.e., the `path` property of an animation channel target that points to such a node **MUST NOT** be any of `"translation"`, `"rotation"`, or `"scale"`. [NOTE] .Rationale ==== Since matrix decomposition is generally ambiguous, animating only one property could result in non-portable implementation-dependent results. Note that `matrix` property presence does not affect the `"weights"` path. ==== Nodes that do not contain a mesh with morph targets **MUST NOT** be targeted with `"weights"` path. Within one animation, each target (a combination of a node and a path) **MUST NOT** be used more than once. [NOTE] .Implementation Note ==== This prevents potential ambiguities when one target is affected by two or more overlapping samplers. ==== Each of the animation's *samplers* defines the `input`/`output` pair: a set of floating-point scalar values representing linear time in seconds; and a set of vectors or scalars representing the animated property. All values are stored in a buffer and accessed via accessors; refer to the table below for output accessor types. Interpolation between keys is performed using the interpolation method specified in the `interpolation` property. Supported `interpolation` values include `LINEAR`, `STEP`, and `CUBICSPLINE`. See <> for additional information about interpolation modes. The inputs of each sampler are relative to `t = 0`, defined as the beginning of the parent `animations` entry. Before and after the provided input range, output **MUST** be clamped to the nearest end of the input range. [NOTE] .Implementation Note ==== For example, if the earliest sampler input for an animation is `t = 10`, a client implementation must begin playback of that animation channel at `t = 0` with output clamped to the first available output value. ==== Samplers within a given animation **MAY** have different inputs. [options="header",cols="15%,15%,35%,35%"] |==== | `channel.path` | Accessor Type | Component Type(s) | Description | `"translation"` | `"VEC3"` | _float32_ | XYZ translation vector | `"rotation"` | `"VEC4"` | _float32_ + _snorm8_ + _unorm8_ + _snorm16_ + _unorm16_ | XYZW rotation quaternion | `"scale"` | `"VEC3"` | _float32_ | XYZ scale vector | `"weights"` | `"SCALAR"` | _float32_ + _snorm8_ + _unorm8_ + _snorm16_ + _unorm16_ | Weights of morph targets |==== Animation sampler's `input` accessor **MUST** have its `min` and `max` properties defined. Animation sampler data accessed with floating-point accessors **MUST NOT** contain infinite or NaN values. [NOTE] .Implementation Note ==== Animations with non-linear time inputs, such as time warps in Autodesk 3ds Max or Maya, are not directly representable with glTF animations. glTF is a runtime format and non-linear time inputs are expensive to compute at runtime. Exporter implementations should sample a non-linear time animation into linear inputs and outputs for an accurate representation. ==== A morph target animation frame is defined by a sequence of scalars of length equal to the number of targets in the animated morph target. These scalar sequences **MUST** lie end-to-end as a single stream in the output accessor, whose final size is equal to the number of morph targets times the number of animation frames. Morph target animation is by nature sparse, consider using <> for storage of morph target animation. When used with `CUBICSPLINE` interpolation, tangents (a~k~, b~k~) and values (v~k~) are grouped within keyframes: a~1~,a~2~,...a~n~,v~1~,v~2~,...v~n~,b~1~,b~2~,...b~n~ See <> for additional information about interpolation modes. Skinned animation is achieved by animating the joints in the skin's joint hierarchy. [[specifying-extensions]] == Specifying Extensions glTF defines an extension mechanism that allows the base format to be extended with new capabilities. Any glTF object **MAY** have an optional `extensions` property, as in the following example: [source,json] ---- { "material": [ { "extensions": { "KHR_materials_sheen": { "sheenColorFactor": [ 1.0, 0.329, 0.1 ], "sheenRoughnessFactor": 0.8 } } } ] } ---- All extensions used in a glTF asset **MUST** be listed in the top-level `extensionsUsed` array object, e.g., [source,json] ---- { "extensionsUsed": [ "KHR_materials_sheen", "VENDOR_physics" ] } ---- All glTF extensions required to load and/or render an asset **MUST** be listed in the top-level `extensionsRequired` array, e.g., [source,json] ---- { "extensionsRequired": [ "KHR_texture_transform" ], "extensionsUsed": [ "KHR_texture_transform" ] } ---- `extensionsRequired` is a subset of `extensionsUsed`. All values in `extensionsRequired` **MUST** also exist in `extensionsUsed`. [[glb-file-format-specification]] = GLB File Format Specification [[glb-file-format-specification-general]] == General (Informative) glTF provides two delivery options that can be used together: * glTF JSON points to external binary data (geometry, key frames, skins), and images. * glTF JSON embeds base64-encoded binary data, and images inline using data URIs. Hence, loading glTF files usually requires either separate requests to fetch all binary data, or extra space due to base64-encoding. Base64-encoding requires extra processing to decode and increases the file size (by ~33% for encoded resources). While transport-layer gzip mitigates the file size increase, decompression and decoding still add significant loading time. To avoid this file size and processing overhead, a container format, _Binary glTF_ is introduced that enables a glTF asset, including JSON, buffers, and images, to be stored in a single binary blob. A Binary glTF asset can still refer to external resources. For example, an application that wants to keep images as separate files may embed everything needed for a scene, except images, in a Binary glTF. [[glb-file-format-specification-structure]] == Structure A Binary glTF (which can be a file, for example) has the following structure: * A 12-byte preamble, called the _header_. * One or more _chunks_ that contain JSON content and binary data. The _chunk_ containing JSON **MAY** refer to external resources as usual, and **MAY** also reference resources stored within other _chunks_. [[glb-file-format-specification-file-extension]] == File Extension & Media Type The file extension to be used with Binary glTF is `.glb`. The registered media type is `model/gltf-binary`. [[binary-gltf-layout]] == Binary glTF Layout [[binary-gltf-layout-overview]] === Overview Binary glTF is little endian. The figure below shows an example of a Binary glTF asset. .Binary glTF Layout image::figures/glb2.svg[pdfwidth=6in,align=left] The following sections describe the structure more in detail. [[binary-header]] === Header The 12-byte header consists of three 4-byte entries: [source,c] ---- uint32 magic uint32 version uint32 length ---- * `magic` **MUST** be equal to equal `0x46546C67`. It is ASCII string `glTF` and can be used to identify data as Binary glTF. * `version` indicates the version of the Binary glTF container format. This specification defines version 2. + Client implementations that load GLB format **MUST** also check for the <> in the JSON chunk, as the version specified in the GLB header only refers to the GLB container version. * `length` is the total length of the Binary glTF, including _header_ and all _chunks_, in bytes. [[chunks]] === Chunks [[chunks-overview]] ==== Overview Each chunk has the following structure: [source,c] ---- uint32 chunkLength uint32 chunkType ubyte[] chunkData ---- * `chunkLength` is the length of `chunkData`, in bytes. * `chunkType` indicates the type of chunk. See <> for details. * `chunkData` is the binary payload of the chunk. The start and the end of each chunk **MUST** be aligned to a 4-byte boundary. See chunks definitions for padding schemes. Chunks **MUST** appear in exactly the order given in <>. [[table-chunktypes]] .Chunk types [options="header"] |==== | | Chunk Type | ASCII | Description | Occurrences | 1. | 0x4E4F534A | JSON | Structured JSON content | 1 | 2. | 0x004E4942 | BIN | Binary buffer | 0 or 1 |==== Client implementations **MUST** ignore chunks with unknown types to enable glTF extensions to reference additional chunks with new types following the first two chunks. [[structured-json-content]] ==== Structured JSON Content This chunk holds the glTF JSON, as it would be provided within a .gltf file. [NOTE] .ECMAScript Implementation Note ==== In a JavaScript implementation, the `TextDecoder` API can be used to extract the glTF content from the ArrayBuffer, and then the JSON can be parsed with `JSON.parse` as usual. ==== This chunk **MUST** be the very first chunk of a Binary glTF asset. By reading this chunk first, an implementation is able to progressively retrieve resources from subsequent chunks. This way, it is also possible to read only a selected subset of resources from a Binary glTF asset. This chunk **MUST** be padded with trailing `Space` chars (`0x20`) to satisfy alignment requirements. [[binary-buffer]] ==== Binary buffer This chunk contains the binary payload for geometry, animation key frames, skins, and images. See <> for details on referencing this chunk from JSON. This chunk **MUST** be the second chunk of the Binary glTF asset. This chunk **MUST** be padded with trailing zeros (`0x00`) to satisfy alignment requirements. When the binary buffer is empty or when it is stored by other means, this chunk **SHOULD** be omitted. [[properties-reference]] = Properties Reference ifndef::revdate[] [NOTE] .Note ==== This section is auto-generated, and will be a broken link when previewing. View the full https://www.khronos.org/registry/glTF/specs/2.0/glTF-2.0.html#properties-reference[Properties Reference Section]. To suggest changes, open a Pull Request to edit the https://github.com/KhronosGroup/glTF/tree/main/specification/2.0/schema[JSON Schema Files]. ==== endif::[] // Generated with wetzel include::PropertiesReference.adoc[] [[acknowledgments]] = Acknowledgments (Informative) [[editors]] == Editors * Saurabh Bhatia, Microsoft * Patrick Cozzi, Cesium * Alexey Knyazev, Individual Contributor * Tony Parisi, Unity [[working-group-and-alumni]] == Khronos 3D Formats Working Group and Alumni * Remi Arnaud, Vario * Mike Bond, Adobe * Leonard Daly, Individual Contributor * Emiliano Gambaretto, Adobe * Tobias Häußler, Dassault Systèmes * Gary Hsu, Microsoft * Marco Hutter, Individual Contributor * Uli Klumpp, Individual Contributor * Max Limper, Fraunhofer IGD * Ed Mackey, Analytical Graphics, Inc. * Don McCurdy, Google * Scott Nagy, Microsoft * Norbert Nopper, UX3D * Fabrice Robinet, Individual Contributor (Previous Editor and Incubator) * Bastian Sdorra, Dassault Systèmes * Neil Trevett, NVIDIA * Jan Paul Van Waveren, Oculus * Amanda Watson, Oculus [[special-thanks]] == Special Thanks * Sarah Chow, Cesium * Tom Fili, Cesium * Darryl Gough * Eric Haines, Autodesk * Yu Chen Hou, Individual Contributor * Scott Hunter, Analytical Graphics, Inc. * Brandon Jones, Google * Arseny Kapoulkine, Individual Contributor * Jon Leech, Individual Contributor * Sean Lilley, Cesium * Juan Linietsky, Godot Engine * Matthew McMullan, Individual Contributor * Mohamad Moneimne, University of Pennsylvania * Kai Ninomiya, formerly Cesium * Cedric Pinson, Sketchfab * Jeff Russell, Marmoset * Miguel Sousa, Fraunhofer IGD * Timo Sturm, Fraunhofer IGD * Rob Taglang, Cesium * Maik Thöner, Fraunhofer IGD * Steven Vergenz, AltspaceVR * Corentin Wallez, Google * Alex Wood, Analytical Graphics, Inc [appendix] [[appendix-a-json-schema-reference]] = JSON Schema Reference (Informative) ifndef::revdate[] [NOTE] .Note ==== This section is auto-generated, and will be a broken link when previewing. View the full https://www.khronos.org/registry/glTF/specs/2.0/glTF-2.0.html#appendix-a-json-schema-reference[JSON Schema Reference]. To suggest changes, open a Pull Request to edit the https://github.com/KhronosGroup/glTF/tree/main/specification/2.0/schema[JSON Schema Files]. ==== endif::[] // Generated with wetzel include::JsonSchemaReference.adoc[] [appendix] [[appendix-b-brdf-implementation]] = BRDF Implementation [[appendix-b-brdf-implementation-general]] == General This chapter presents the bidirectional reflectance distribution function (BRDF) of the glTF 2.0 metallic-roughness material. The BRDF describes the reflective properties of the surface of a physically based material. For a pair of directions, the BRDF returns how much light from the incoming direction is reflected from the surface in the outgoing direction. [NOTE] .Note ==== See <>, for an introduction to radiometry and the BRDF. ==== The BRDF of the metallic-roughness material is a linear interpolation of a metallic BRDF and a dielectric BRDF. The BRDFs share the parameters for roughness and base color. The blending factor `metallic` describes the metalness of the material. [source,c] ---- material = mix(dielectric_brdf, metal_brdf, metallic) = (1.0 - metallic) * dielectric_brdf + metallic * metal_brdf ---- [NOTE] .Note ==== Such a material model based on a linear interpolation of metallic and dielectric components was introduced by <> and adapted by many renderers, resulting in a wide range of applications supporting it. Usually, a material is either metallic or dielectric. A texture provided for `metallic` with either `1.0` or `0.0` separates metallic from dielectric regions on the mesh. There are situations in which there is no clear separation. It may happen that due to anti-aliasing or mip-mapping there is a portion of metal and a portion of dielectric within a texel. Furthermore, a material composed of several semi-transparent layers may be represented as a blend between several single-layered materials (layering via parameter blending). ==== The logical structure of the material is presented below, using an abstract notation that describes the material as a directed acyclic graph (DAG). The vertices correspond to the basic building blocks of the material model: BRDFs, mixing operators, input parameters, and constants. This is followed by an informative sample implementation as a set of equations and source code for the BRDFs and mixing operators. [[material-structure]] == Material Structure [[metals]] === Metals Metallic surfaces reflect most illumination, only a small portion of the light is absorbed by the material. [NOTE] .Note ==== See <>. ==== This effect is described by the Fresnel term `conductor_fresnel` with the wavelength-dependent refractive index and extinction coefficient. To make parameterization simple, the metallic-roughness material combines the two quantities into a single, user-defined color value `baseColor` that defines the reflection color at normal incidence, also referred to as `f0`. The reflection color at grazing incidence is called `f90`. It is set to `1.0` because the grazing angle reflectance for any material approaches pure white in the limit. The conductor Fresnel term modulates the contribution of a specular BRDF parameterized by the `roughness` parameter. [source,c] ---- metal_brdf = conductor_fresnel( f0 = baseColor, bsdf = specular_brdf( α = roughness ^ 2)) ---- [[dielectrics]] === Dielectrics Unlike metals, dielectric materials transmit most of the incident illumination into the interior of the object and the Fresnel term is parameterized only by the refractive index. [NOTE] .Note ==== See <>. ==== This makes dielectrics like glass, oil, water or air transparent. Other dielectrics, like most plastic materials, are filled with particles that absorb or scatter most or all of the transmitted light, reducing the transparency and giving the surface its colorful appearance. As a result, dielectric materials are modeled as a Fresnel-weighted combination of a specular BRDF, simulating the reflection at the surface, and a diffuse BRDF, simulating the transmitted portion of the light that is absorbed and scattered inside the object. The reflection roughness is given by the squared `roughness` of the material. The color of the diffuse BRDF comes from the `baseColor`. The amount of reflection compared to transmission is directional-dependent and as such determined by the Fresnel term. Its index of refraction is set to a fixed value of `1.5`, a good compromise for most opaque, dielectric materials. [source,c] ---- dielectric_brdf = fresnel_mix( ior = 1.5, base = diffuse_brdf( color = baseColor), layer = specular_brdf( α = roughness ^ 2)) ---- [[microfacet-surfaces]] === Microfacet Surfaces The metal BRDF and the dielectric BRDF are based on a microfacet model. [NOTE] .Note ==== The theory behind microfacet models was developed in early works by <>, <> and others. ==== A microfacet model describes the orientation of tiny facets (microfacets) on the surface as a statistical distribution. The distribution determines the orientation of the facets as a random perturbation around the normal direction of the surface. The perturbation strength depends on the `roughness` parameter and varies between `0.0` (smooth surface) and `1.0` (rough surface). A number of distribution functions have been proposed in the last decades. The Trowbridge-Reitz / GGX microfacet distribution describes the microsurface as being composed of perfectly specular, infinitesimal oblate ellipsoids, whose half-height in the normal direction is α times the radius in the tangent plane. α = 1 gives spheres, which results in uniform reflection in all directions. This reflection behavior corresponds to a rough surface. α = 0 gives a perfectly specular surface. [NOTE] .Note ==== The Trowbridge-Reitz distribution was first described by <>. Later <> independently developed the same distribution and called it "`GGX`". They show that it is a better fit for measured data than the Beckmann distribution used by <> due to its stronger tails. ==== The mapping α = `roughness`^2^ results in more perceptually linear changes in the roughness. [NOTE] .Note ==== This mapping was suggested by <>. ==== The distribution only describes the proportion of each normal on the microsurface. It does not describe how the normals are organized. For this we need a microsurface profile. [NOTE] .Note ==== The difference between distribution and profile is detailed by <>, where he in addition provides an extensive study of common microfacet profiles. Based on this work, we suggest using the Smith microsurface profile (originally developed by <>) and its corresponding masking-shadowing function. Heitz describes the Smith profile as the most accurate model for reflection from random height fields. It assumes that height and normal between neighboring points are not correlated, implying a random set of microfacets instead of a continuous surface. Microfacet models often do not consider multiple scattering. The shadowing term suppresses light that intersects the microsurface a second time. <> extended the Smith-based microfacet models to include a multiple scattering component, which significantly improves accuracy of predictions of the model. We suggest to incorporate multiple scattering whenever possible, either by making use of the unbiased stochastic evaluation introduced by Heitz, or one of the approximations presented later, for example by <> or <>. ==== [[complete-model]] === Complete Model The BRDFs and mixing operators used in the metallic-roughness material are summarized in the following figure. .BRDFs and Mixing Operators image::figures/pbr.svg[pdfwidth=4in,align=left] The glTF Specification is designed to allow applications to choose different lighting implementations based on their requirements. Some implementations **MAY** focus on an accurate simulation of light transport while others **MAY** choose to deliver real-time performance. Therefore, any implementation that adheres to the rules for mixing BRDFs is conformant to the glTF Specification. In a physically accurate light simulation, the BRDFs **MUST** follow some basic principles: the BRDF **MUST** be positive, reciprocal, and energy conserving. This ensures that the visual output of the simulation is independent of the underlying rendering algorithm, if it is unbiased. The unbiased light simulation with physically realistic BRDFs will be the ground-truth for approximations in real-time renderers that are often biased, but still give visually pleasing results. [[implementation]] == Sample Implementation (Informative) [[implementation-overview]] === Overview Often, renderers use approximations to solve the rendering equation, like the split-sum approximation for image based lighting, or simplify the math to save instructions and reduce register pressure. However, there are many ways to achieve good approximations, depending on the platform. A sample implementation is available at https://github.com/KhronosGroup/glTF-Sample-Viewer/ and provides an example of a WebGL 2.0 implementation of a standard BRDF based on the glTF material parameters. To achieve high performance in real-time applications, this implementation uses several approximations and uses non-physical simplifications that break energy-conservation and reciprocity. We use the following notation: * *V* is the normalized vector from the shading location to the eye * *L* is the normalized vector from the shading location to the light * *N* is the surface normal in the same space as the above values * *H* is the half vector, where *H* = normalize(*L* + *V*) [[specular-brdf]] === Specular BRDF ifndef::revdate[] [NOTE] .Note ==== This section contains formulas that may not display correctly in all preview tools. View the full https://www.khronos.org/registry/glTF/specs/2.0/glTF-2.0.html#specular-brdf[Specular BRDF Section]. ==== endif::[] The specular reflection `specular_brdf(α)` is a microfacet BRDF [latexmath] ++++ \text{MicrofacetBRDF} = \frac{G D}{4 \, \left|N \cdot L \right| \, \left| N \cdot V \right|} ++++ with the Trowbridge-Reitz/GGX microfacet distribution [latexmath] ++++ D = \frac{\alpha^2 \, \chi^{+}(N \cdot H)}{\pi ((N \cdot H)^2 (\alpha^2 - 1) + 1)^2} ++++ and the height-correlated form of the Smith joint masking-shadowing function [latexmath] ++++ G = \frac{2 \, \left| N \cdot L \right| \, \left| N \cdot V \right| \, \chi^{+}(H \cdot L) \, \chi^{+}(H \cdot V)}{\left| N \cdot V \right| \, \sqrt{\alpha^2 + (1 - \alpha^2) (N \cdot L)^2} + \left| N \cdot L \right| \, \sqrt{\alpha^2 + (1 - \alpha^2) (N \cdot V)^2}} ++++ where χ^+^(*x*) denotes the Heaviside function: 1 if *x* > 0 and 0 if *x* <= 0. See <> for a derivation of the formulas. Introducing the visibility function latexmath:[\nu] [latexmath] ++++ \nu = \frac{G}{4 \, \left| N \cdot L \right| \, \left| N \cdot V \right|} ++++ simplifies the original microfacet BRDF to [latexmath] ++++ \text{MicrofacetBRDF} = \nu D ++++ with [latexmath] ++++ \nu = \frac{\chi^{+}(H \cdot L) \, \chi^{+}(H \cdot V)}{2 \, (\left| N \cdot V \right| \, \sqrt{\alpha^2 + (1 - \alpha^2) (N \cdot L)^2} + \left| N \cdot L \right| \, \sqrt{\alpha^2 + (1 - \alpha^2) (N \cdot V)^2})} ++++ Thus, we have the function [source,c] ---- function specular_brdf(α) { return Vis * D } ---- [NOTE] ==== A roughness of zero (α = 0) cannot be evaluated directly with this formulation. As α approaches zero, the GGX distribution latexmath:[D] collapses to a delta function and the BRDF becomes singular, leading to divisions by zero or numerically unstable results. Implementations should therefore never use α = 0 in the equations above. Instead, α should be clamped to a small positive value, or the surface should be handled as an ideal specular (mirror) reflector using a separate code path. ==== [[diffuse-brdf]] === Diffuse BRDF ifndef::revdate[] [NOTE] .Note ==== This section contains formulas that may not display correctly in all preview tools. View the full https://www.khronos.org/registry/glTF/specs/2.0/glTF-2.0.html#diffuse-brdf[Diffuse BRDF Section]. ==== endif::[] The diffuse reflection `diffuse_brdf(color)` is a Lambertian BRDF [latexmath] ++++ \text{LambertianBRDF} = \frac{1}{\pi} ++++ multiplied with the `color`. [source,c] ---- function diffuse_brdf(color) { return (1/pi) * color } ---- [[fresnel]] === Fresnel ifndef::revdate[] [NOTE] .Note ==== This section contains formulas that may not display correctly in all preview tools. View the full https://www.khronos.org/registry/glTF/specs/2.0/glTF-2.0.html#fresnel[Fresnel Section]. ==== endif::[] An inexpensive approximation for the Fresnel term that can be used for conductors and dielectrics was developed by <>: [latexmath] ++++ F = f_0 + (1 - f_0) (1 - \left| V \cdot H \right| )^5 ++++ The conductor Fresnel `conductor_fresnel(f0, bsdf)` applies a view-dependent tint to a BSDF: [source,c] ---- function conductor_fresnel(f0, bsdf) { return bsdf * (f0 + (1 - f0) * (1 - abs(VdotH))^5) } ---- For the dielectric BRDF a diffuse component `base` and a specular component `layer` are combined via `fresnel_mix(ior, base, layer)`. The `f0` color is now derived from the index of refraction `ior`. [source,c] ---- function fresnel_mix(ior, base, layer) { f0 = ((1-ior)/(1+ior))^2 fr = f0 + (1 - f0)*(1 - abs(VdotH))^5 return mix(base, layer, fr) } ---- [[metal-brdf-and-dielectric-brdf]] === Metal BRDF and Dielectric BRDF Now that we have an implementation for all the functions used in the glTF metallic-roughness material model, we are able to connect the functions according to the graph shown in section <>. By substituting the mixing functions (`fresnel_mix`, `conductor_fresnel`) for the implementation, we arrive at the following BRDFs for the metal and the dielectric component: [source,c] ---- metal_brdf = specular_brdf(roughness^2) * (baseColor.rgb + (1 - baseColor.rgb) * (1 - abs(VdotH))^5) dielectric_brdf = mix(diffuse_brdf(baseColor.rgb), specular_brdf(roughness^2), 0.04 + (1 - 0.04) * (1 - abs(VdotH))^5) ---- Note that the dielectric index of refraction `ior = 1.5` is now `f0 = 0.04`. Metal and dielectric are mixed according to the metalness: [source,c] ---- material = mix(dielectric_brdf, metal_brdf, metallic) ---- The full code is given below: [source,c] ---- fresnel_w = (1 - abs(VdotH))^5 diffuse_brdf = (1 / π) * baseColor.rgb specular_brdf = D(roughness^2) * G(roughness^2) / (4 * abs(VdotN) * abs(LdotN)) dielectric_f0 = 0.04 dielectric_fresnel = dielectric_f0 + (1 - dielectric_f0) * fresnel_w dielectric_brdf = mix(diffuse_brdf, specular_brdf, dielectric_fresnel) metal_fresnel = baseColor.rgb + (1 - baseColor.rgb) * fresnel_w metal_brdf = metal_fresnel * specular_brdf material = mix(dielectric_brdf, metal_brdf, metallic) ---- [[discussion]] === Discussion [[masking-shadowing-term-and-multiple-scattering]] ==== Masking-Shadowing Term and Multiple Scattering The model for specular reflection can be improved in several ways. <> notes that a more accurate form of the masking-shadowing function takes the correlation between masking and shadowing due to the height of the microsurface into account. This correlation is accounted for in the height-correlated masking and shadowing function. Another improvement in accuracy can be achieved by modeling multiple scattering, see Section <>. [[schlicks-fresnel-approximation]] ==== Schlick's Fresnel Approximation Although Schlick's Fresnel is a good approximation for a wide range of metallic and dielectric materials, there are a couple of reasons to use a more sophisticated solution for the Fresnel term. Metals often exhibit a "`dip`" in reflectance near grazing angles, which is not present in the Schlick Fresnel. <> extend the Schlick Fresnel with an error term to account for it. <> improves the parameterization of this term by introducing an artist-friendly `f82` color, the color at an angle of about 82°. An additional color parameter for metals was also introduced by <>. Gulbrandson calls it "`edge tint`" and uses it in the full Fresnel equations instead of Schlick's approximation. Even though the full Fresnel equations should give a more accurate result, Hoffman shows that it is worse than Schlick's approximation in the context of RGB renderers. As we target RGB renderers and do not provide an additional color parameter for metals in glTF, we suggest using the original Schlick Fresnel for metals. The index of refraction of most dielectrics is 1.5. For that reason, the dielectric Fresnel term uses a fixed `f0 = 0.04`. The Schlick Fresnel approximates the full Fresnel equations well for an index of refraction in the range [1.2, 2.2]. The main reason for a material to fall outside this range is transparency and nested objects. If a transparent object overlaps another transparent object and both have the same (or similar) index of refraction, the resulting ratio at the boundary is 1 (or close to 1). According to the full Fresnel equations, there is no (or almost no) reflection in this case. The reflection intensity computed from the Schlick Fresnel approximation will be too high. Implementations that care about accuracy in case of nested dielectrics are encouraged to use the full Fresnel equations for dielectrics. For metals Schlick's approximation is still a good choice. [[coupling-diffuse-and-specular-reflection]] ==== Coupling Diffuse and Specular Reflection While the coupling of diffuse and specular components in `fresnel_mix` as proposed in this section is simple and cheap to compute, it is not very accurate and breaks a fundamental property that a physically based BRDF must fulfill — energy conservation. Energy conservation means that a BRDF must not reflect more light than it receives. Several fixes have been proposed, each with its own trade-offs regarding performance and quality. <> notes that a common solution found in many models calculates the diffuse Fresnel factor by evaluating the Fresnel term twice with view and light direction instead of the half vector: `(1-F(NdotL)) * (1-F(NdotV))`. While this is energy-conserving, he notes that this weighting results in significant darkening at grazing angles, an effect they couldn't observe in their measurements. They propose some changes to the diffuse BRDF to make it better predict the measurements, but even the fixed version is still not energy conserving mathematically. More recently, <> developed a generic framework for computing BSDFs of layered materials, including multiple scattering within layers. Amongst much more complicated scenarios it also solves the special case of coupling diffuse and specular components, but it is too heavy for textured materials, even in offline rendering. <> found a solution tailored to the special case of coupling diffuse and specular components, which is easy to compute. It requires the directional albedo of the Fresnel-weighted specular BRDF to be precomputed and tabulated, but they found that the function is smooth, and a low-resolution 3D texture (16³ pixels) is sufficient. Their coupled diffuse-specular model is not only energy-__con__serving, but also energy-__pre__serving, meaning that if neither the specular nor the diffuse component absorb any energy, all energy is reflected. [[references]] == References * [[Burley2012]] https://disneyanimation.com/publications/physically-based-shading-at-disney/[Burley, B. (2012): Physically-Based Shading at Disney.] * [[CookTorrance1982]] https://graphics.pixar.com/library/ReflectanceModel/paper.pdf[Cook, R. L., and K. E. Torrance (1982): A Reflectance Model for Computer Graphics. ACM Transactions on Graphics 1 (1), 7-24.] * [[Gulbrandsen2014]] http://jcgt.org/published/0003/04/03/paper-lowres.pdf[Gulbrandsen, O. (2014): Artist Friendly Metallic Fresnel] * [[Heitz2014]] http://jcgt.org/published/0003/02/03/paper.pdf[Heitz, E. (2014): Understanding the Masking-Shadowing Function in Microfacet-Based BRDFs] * [[Heitz2016]] https://eheitzresearch.wordpress.com/240-2/[Heitz, E., J. Hanika, E. d'Eon, and C. Dachsbacher (2016): Multiple-Scattering Microfacet BSDFs with the Smith Model] * [[Hoffman2019]] https://renderwonk.com/publications/mam2019/[Naty Hoffman (2019): Fresnel Equations Considered Harmful] * [[Jakob2014]] https://research.cs.cornell.edu/layered-sg14/[Jakob, W., E. d'Eon, O. Jakob, S. Marschner (2014): A Comprehensive Framework for Rendering Layered Materials] * [[KullaConty2017]] https://blog.selfshadow.com/publications/s2017-shading-course/imageworks/s2017_pbs_imageworks_slides_v2.pdf[Kulla, C., and A. Conty (2017): Revisiting Physically Based Shading at Imageworks] * [[LazanyiSzirmayKalos2005]] http://wscg.zcu.cz/WSCG2005/Papers_2005/Short/H29-full.pdf[Lazanyi, I. and L. Szirmay-Kalos (2005): Fresnel term approximations for metals] * [[Pharr2016]] https://www.pbr-book.org/[Pharr, M., W. Jakob, and G. Humphreys (2016): Physically Based Rendering: From Theory To Implementation, 3rd edition.] * [[Schlick1994]] https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.50.2297&rep=rep1&type=pdf[Schlick, C. (1994): An Inexpensive BRDF Model for Physically-based Rendering. Computer Graphics Forum 13, 233-246.] * [[Smith1967]] https://ieeexplore.ieee.org/document/1138991[Smith, B. (1967): Geometrical shadowing of a random rough surface. IEEE Transactions on Antennas and Propagation 15 (5), 668-671.] * [[TorranceSparrow1967]] https://www.graphics.cornell.edu/~westin/pubs/TorranceSparrowJOSA1967.pdf[Torrance, K. E., E. M. Sparrow (1967): Theory for Off-Specular Reflection From Roughened Surfaces. Journal of the Optical Society of America 57 (9), 1105-1114.] * [[TrowbridgeReitz1975]] https://www.osapublishing.org/josa/abstract.cfm?uri=josa-65-5-531[Trowbridge, T., and K. P. Reitz (1975): Average irregularity representation of a rough surface for ray reflection. Journal of the Optical Society of America 65 (5), 531-536.] * [[Turquin2019]] https://blog.selfshadow.com/publications/turquin/ms_comp_final.pdf[Turquin E. (2019): Practical multiple scattering compensation for microfacet models] * [[Walter2007]] https://www.cs.cornell.edu/~srm/publications/EGSR07-btdf.html[Walter, B., S. Marschner, H. Li, and K. Torrance (2007): Microfacet models for refraction through rough surfaces.] [appendix] [[appendix-c-interpolation]] = Animation Sampler Interpolation Modes == Overview ifndef::revdate[] [NOTE] .Note ==== This appendix contains formulas that may not display correctly in all preview tools. View the full https://www.khronos.org/registry/glTF/specs/2.0/glTF-2.0.html#appendix-c-interpolation[Appendix C: Animation Sampler Interpolation Modes]. ==== endif::[] Animation sampler interpolation modes define how to compute values of animated properties for the timestamps located between the keyframes. When the current (requested) timestamp exists in the animation data, its associated property value is used as-is, without interpolation. For the following sections, let [none] * stem:[n] be the total number of keyframes, stem:[n > 0]; * stem:[t_k] be the timestamp of the stem:[k]-th keyframe, stem:[k \in [1,n\]]; * stem:[v_k] be the animated property value of the stem:[k]-th keyframe; * stem:[t_c] be the current (requested) timestamp, stem:[t_k < t_c < t_{k+1}]; * stem:[t_d = t_{k + 1} - t_k] be the duration of the interpolation segment; * stem:[t = (t_c - t_k) / t_d] be the segment-normalized interpolation factor. The scalar-vector multiplications are per vector component. == Step Interpolation This mode is used when the animation sampler interpolation mode is set to `STEP`. The interpolated sampler value stem:[v_t] at the current (requested) timestamp stem:[t_c] is computed as follows. [latexmath] ++++ v_t = v_k ++++ [[interpolation-lerp]] == Linear Interpolation This mode is used when the animation sampler interpolation mode is set to `LINEAR` and the animated property is not `rotation`. The interpolated sampler value stem:[v_t] at the current (requested) timestamp stem:[t_c] is computed as follows. [latexmath] ++++ v_t = (1 - t) \cdot v_k + t \cdot v_{k+1} ++++ [[interpolation-slerp]] == Spherical Linear Interpolation This mode is used when the animation sampler interpolation mode is set to `LINEAR` and the animated property is `rotation`, i.e., values of the animated property are unit quaternions. Let [none] * latexmath:[a = \arccos(|v_k \cdot v_{k+1}|)] be the arccosine of the absolute value of the dot product of two consecutive quaternions; * latexmath:[s = \frac{v_k \cdot v_{k+1}}{|v_k \cdot v_{k+1}|}] be the sign of the dot product of two consecutive quaternions. The interpolated sampler value stem:[v_t] at the timestamp stem:[t_c] is computed as follows. [latexmath] ++++ v_t = \frac{\sin(a \cdot (1 - t))}{\sin(a)} \cdot v_k + s \cdot \frac{\sin(a \cdot t)}{\sin(a)} \cdot v_{k+1} ++++ [NOTE] .Implementation Note ==== Using the dot product's absolute value for computing stem:[a] and multiplying stem:[v_{k+1}] by the dot product's sign ensure that the spherical interpolation follows the short path along the great circle defined by the two quaternions. ==== Implementations **MAY** approximate these equations to reach application-specific accuracy and/or performance targets. [NOTE] .Implementation Note ==== When stem:[a] is close to zero, spherical linear interpolation turns into regular linear interpolation. ==== [[interpolation-cubic]] == Cubic Spline Interpolation This mode is used when the animation sampler interpolation mode is set to `CUBICSPLINE`. An animation sampler that uses cubic spline interpolation **MUST** have at least 2 keyframes. For each timestamp stored in the animation sampler, there are three associated keyframe values: in-tangent, property value, and out-tangent. Let [none] * stem:[a_k], stem:[v_k], and stem:[b_k] be the in-tangent, the property value, and the out-tangent of the stem:[k]-th frame respectively. The interpolated sampler value stem:[v_t] at the timestamp stem:[t_c] is computed as follows. [latexmath] ++++ v_t = (2t^3 - 3t^2 + 1) \cdot v_k + t_d \cdot (t^3 - 2t^2 + t) \cdot b_k + (-2t^3 + 3t^2) \cdot v_{k+1} + t_d \cdot (t^3 - t^2) \cdot a_{k+1} ++++ When the animation sampler targets a node's rotation property, the interpolated quaternion **MUST** be normalized before applying the result to the node's rotation. When writing out rotation values, exporters **SHOULD** take care to not write out values that can result in an invalid quaternion with all zero values being produced by the interpolation. [NOTE] .Implementation Note ==== This can be achieved by ensuring that stem:[v_k != -v_{k+1}] for all keyframes. ==== The first in-tangent stem:[a_1] and last out-tangent stem:[b_n] **SHOULD** be zeros as they are not used in the spline calculations.