# File Storage Extension The File Storage Extension can persist state to the local file system. | Status | | | ------------- |-----------| | Stability | [beta] | | Distributions | [contrib], [k8s] | | Issues | [![Open issues](https://img.shields.io/github/issues-search/open-telemetry/opentelemetry-collector-contrib?query=is%3Aissue%20is%3Aopen%20label%3Aextension%2Ffilestorage%20&label=open&color=orange&logo=opentelemetry)](https://github.com/open-telemetry/opentelemetry-collector-contrib/issues?q=is%3Aopen+is%3Aissue+label%3Aextension%2Ffilestorage) [![Closed issues](https://img.shields.io/github/issues-search/open-telemetry/opentelemetry-collector-contrib?query=is%3Aissue%20is%3Aclosed%20label%3Aextension%2Ffilestorage%20&label=closed&color=blue&logo=opentelemetry)](https://github.com/open-telemetry/opentelemetry-collector-contrib/issues?q=is%3Aclosed+is%3Aissue+label%3Aextension%2Ffilestorage) | | Code coverage | [![codecov](https://codecov.io/github/open-telemetry/opentelemetry-collector-contrib/graph/main/badge.svg?component=extension_filestorage)](https://app.codecov.io/gh/open-telemetry/opentelemetry-collector-contrib/tree/main/?components%5B0%5D=extension_filestorage&displayType=list) | | [Code Owners](https://github.com/open-telemetry/opentelemetry-collector-contrib/blob/main/CONTRIBUTING.md#becoming-a-code-owner) | [@swiatekm](https://www.github.com/swiatekm), [@VihasMakwana](https://www.github.com/VihasMakwana) \| Seeking more code owners! | | Emeritus | [@djaglowski](https://www.github.com/djaglowski) | [beta]: https://github.com/open-telemetry/opentelemetry-collector/blob/main/docs/component-stability.md#beta [contrib]: https://github.com/open-telemetry/opentelemetry-collector-releases/tree/main/distributions/otelcol-contrib [k8s]: https://github.com/open-telemetry/opentelemetry-collector-releases/tree/main/distributions/otelcol-k8s The File Storage extension can persist state to the local file system. The extension requires read and write access to a directory. A default directory can be used, but it must already exist in order for the extension to operate. `directory` is the relative or absolute path to the dedicated data storage directory. The default directory is `%ProgramData%\Otelcol\FileStorage` on Windows and `/var/lib/otelcol/file_storage` otherwise. `timeout` is the maximum time to wait for a file lock. This value does not need to be modified in most circumstances. The default timeout is `1s`. `max_size` sets the maximum on-disk size of each bbolt database file in bytes. When a write would need the file to grow past this limit, the write is rejected with a storage-full error. A value of `0` means unlimited size. Writes that fit into already-allocated free space are still allowed, even when the file is already at the configured limit. When rebound compaction is enabled, `max_size` must be greater than or equal to both `compaction.rebound_needed_threshold_mib * 1,048,576` and `compaction.rebound_trigger_threshold_mib * 1,048,576`. `fsync` when set, will force the database to perform an fsync after each write. This helps to ensure database integrity if there is an interruption to the database process, but at the cost of performance. See [DB.NoSync](https://pkg.go.dev/go.etcd.io/bbolt#DB) for more information. `create_directory` when set, will create the data storage and compaction directories if they do not already exist. By default, the directories will be created with `0750 (rwxr-x---)` permissions, minus the process umask. Use `directory_permissions` to customize directory creation permissions, minus the process umask. `recreate` when set, the filestorage extension will automatically rename the corrupted bbolt database and create a new one when certain bbolt panics occur. See (#35899)[https://github.com/open-telemetry/opentelemetry-collector-contrib/issues/35899] for more details. If the database fails to open due to corruption (resulting in a panic), the corrupted file will be automatically renamed to `{filename}.{ISO 8601 timestamp}.backup` and a new data file will be created from scratch. This allows the collector to continue operating even when encountering certain bbolt panics. If no corruption is detected, the existing database continues to be used normally. There may still be scenarios where manually removing or renaming the file may be required, and this feature flag is not a panacea for all bbolt panics you can encounter. > [!Note] > When database corruption is detected and automatic recovery is triggered, the corrupted data will be moved to a `.backup` file. While this prevents complete data loss, the collector will start with a fresh database, which may lead to data duplication or loss of component state. ## Compaction `compaction` defines how and when files should be compacted. There are two modes of compaction available (both of which can be set concurrently): - `compaction.on_start` (default: false), which happens when collector starts - `compaction.on_rebound` (default: false), which happens online when certain criteria are met; it's discussed in more detail below `compaction.directory` specifies the directory used for compaction (as a midstep). `compaction.max_transaction_size` (default: 65536): defines maximum size of the compaction transaction. A value of zero will ignore transaction sizes. `compaction.cleanup_on_start` (default: false) - specifies if removal of compaction temporary files is performed on start. It will remove all temporary files in the compaction directory (those which start with `tempdb`), temp files will be left if a previous run of the process is killed while compacting. If `max_size` is set, both `compaction.rebound_needed_threshold_mib` and `compaction.rebound_trigger_threshold_mib` must be less than or equal to that limit after converting MiB to bytes. ### Rebound (online) compaction For rebound compaction, there are two additional parameters available: - `compaction.rebound_needed_threshold_mib` (default: 100) - when allocated data exceeds this amount, the "compaction needed" flag will be enabled - `compaction.rebound_trigger_threshold_mib` (default: 10) - if the "compaction needed" flag is set and allocated data drops below this amount, compaction will begin and the "compaction needed" flag will be cleared - `compaction.check_interval` (default: 5s) - specifies how frequently the conditions for compaction are being checked The idea behind rebound compaction is that in certain workloads (e.g. [persistent queue](https://github.com/open-telemetry/opentelemetry-collector/tree/main/exporter/exporterhelper#persistent-queue)) the storage might grow significantly (e.g. when the exporter is unable to send the data due to network problem) after which it is being emptied as the underlying issue is gone (e.g. network connectivity is back). This leaves a significant space that needs to be reclaimed (also, this space is reported in memory usage as mmap() is used underneath). The optimal conditions for this to happen online is after the storage is largely drained, which is being controlled by `rebound_trigger_threshold_mib`. To make sure this is not too sensitive, there's also `rebound_needed_threshold_mib` which specifies the total claimed space size that must be met for online compaction to even be considered. Consider following diagram for an example of meeting the rebound (online) compaction conditions. ``` ▲ │ │ XX............. m │ XXXX............ e ├───────────XXXXXXX..........──────────── rebound_needed_threshold_mib m │ XXXXXXXXX.......... o │ XXXXXXXXXXX......... r │ XXXXXXXXXXXXXXXXX.... y ├─────XXXXXXXXXXXXXXXXXXXXX..──────────── rebound_trigger_threshold_mib │ XXXXXXXXXXXXXXXXXXXXXXXXXX......... │ XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX └──────────────── time ─────────────────► │ | | issue draining compaction happens starts begins and reclaims space X - actually used space . - claimed but no longer used space ``` ## Example ```yaml extensions: file_storage: file_storage/all_settings: directory: /var/lib/otelcol/mydir timeout: 1s max_size: 268435456 recreate: true compaction: on_start: true directory: /tmp/ max_transaction_size: 65_536 fsync: false service: extensions: [file_storage, file_storage/all_settings] pipelines: traces: receivers: [nop] exporters: [nop] # Data pipeline is required to load the config. receivers: nop: exporters: nop: ``` ## Replacing unsafe characters in component names The extension uses the type and name of the component using the extension to create a file where the component's data is stored. For example, if a Journald receiver named `journald/myservice` uses the extension, its data is stored in a file named `receiver_journald_myservice`. Sometimes the component name contains characters that either have special meaning in paths - like `/` - or are problematic or even forbidden in file names (depending on the host operating system), like `?` or `|`. To prevent surprising or erroneous behavior, some characters in the component names are replaced before creating the file name to store data by the extension. For example, for a Journald receiver named `journal/mynamespace/myservice`, the component name `mynamespace/myservice` is sanitized into `mynamespace~002Fmyservice` and the data is stored in a file named `receiver_journald_mynamespace~002Fmyservice`. Every unsafe character is replaced with a tilde `~` and the character's [Unicode number][unicode_chars] in hex. The only safe characters are: uppercase and lowercase ASCII letters `A-Z` and `a-z`, digits `0-9`, dot `.`, hyphen `-`, underscore `_`. The tilde `~` character is also replaced even though it is a safe character, to make sure that the sanitized component name never overlaps with a component name that does not require sanitization. [unicode_chars]: https://en.wikipedia.org/wiki/List_of_Unicode_characters ## Filename length handling When creating storage files, the extension derives file names from the type and name of the component using the extension (after sanitizing unsafe characters as described above). In some environments, especially when component names are long or deeply nested, the resulting file name may exceed the maximum file name length allowed by the operating system. Previously, this could cause the extension to fail with a “filename too long” error. The extension now automatically truncates generated file names when necessary to ensure they stay within OS limits. Truncation is applied after sanitization and preserves uniqueness to avoid collisions between different components. To determine which truncated file corresponds to a specific component, check the extension logs. The logs include mappings between component identifiers (type and full name) and the generated file names, allowing you to match components to their stored files. ## Troubleshooting _Currently, the File Storage extension uses [bbolt](https://github.com/etcd-io/bbolt) to store and read data on disk. The following troubleshooting method works for bbolt-managed files. As such, there is no guarantee that this method will continue to work in the future, particularly if the extension switches away from bbolt._ When troubleshooting components that use the File Storage extension, it is sometimes helpful to read the raw contents of files created by the extension for the component. The simplest way to read files created by the File Storage extension is to use the strings utility ([Linux](https://man7.org/linux/man-pages/man1/strings.1.html), [Windows](https://learn.microsoft.com/en-us/sysinternals/downloads/strings)). For example, here are the contents of the file created by the File Storage extension when it's configured as the storage for the File Log receiver. ```sh $ strings /tmp/otelcol/file_storage/filelogreceiver/receiver_filelog_ default file_input.knownFiles2 {"Fingerprint":{"first_bytes":"MzEwNzkKMjE5Cg=="},"Offset":10,"FileAttributes":{"log.file.name":"1.log"},"HeaderFinalized":false,"FlushState":{"LastDataChange":"2024-03-20T18:16:18.164331-07:00","LastDataLength":0}} {"Fingerprint":{"first_bytes":"MjQ0MDMK"},"Offset":6,"FileAttributes":{"log.file.name":"2.log"},"HeaderFinalized":false,"FlushState":{"LastDataChange":"2024-03-20T18:16:39.96429-07:00","LastDataLength":0}} default file_input.knownFiles2 {"Fingerprint":{"first_bytes":"MzEwNzkKMjE5Cg=="},"Offset":10,"FileAttributes":{"log.file.name":"1.log"},"HeaderFinalized":false,"FlushState":{"LastDataChange":"2024-03-20T18:16:18.164331-07:00","LastDataLength":0}} {"Fingerprint":{"first_bytes":"MjQ0MDMK"},"Offset":6,"FileAttributes":{"log.file.name":"2.log"},"HeaderFinalized":false,"FlushState":{"LastDataChange":"2024-03-20T18:16:39.96429-07:00","LastDataLength":0}} default file_input.knownFiles2 {"Fingerprint":{"first_bytes":"MzEwNzkKMjE5Cg=="},"Offset":10,"FileAttributes":{"log.file.name":"1.log"},"HeaderFinalized":false,"FlushState":{"LastDataChange":"2024-03-20T18:16:18.164331-07:00","LastDataLength":0}} {"Fingerprint":{"first_bytes":"MjQ0MDMK"},"Offset":6,"FileAttributes":{"log.file.name":"2.log"},"HeaderFinalized":false,"FlushState":{"LastDataChange":"2024-03-20T18:16:39.96429-07:00","LastDataLength":0}} ```