## Upcoming Changes ### Breaking ### Features * timestamper: support generating timestamps from the current time when no `source_fields` are configured * field_manager: add flag for deactivating deduplication * calculator: extended expression functionality ### Improvements * timestamper: validate `source_format` and `source_timezone` according to the configured `source_fields` * calculator: optimized runtime expression evaluation * docs: improve documentation around dynamic templating for `generic_adder` ### Bugfix * decoder: corrected rfc 5324 in docs to 5424 * field_manager: allow for copying fields only containing `false` and `0` * key_checker: remove configuration fields that were inherited but didnt do anything ## 21.0.0 ### Breaking * restrict UNIX timestamp normalization to seconds, milliseconds, microseconds, and nanoseconds. * handle non-string `concatenator` source fields with `ProcessingWarning` instead of `ProcessingCriticalError` * filter: field values are now automatically coerced in range matches ### Features * add support for fractional UNIX timestamps in the `timestamper` processor while preserving supported integer timestamp normalization. * introduce API-level support for asynchronous rule processing and I/O capability detection in `ng` processors * generic_adder: add support for templated http urls & content_field * generic_resolver: add content_field support * field_name_replacer: add new `field_name_replacer` processor to replace occurences of strings in key names * filter: add support for `*` (open boundary) in range expressions * filter: allow mixed numeric range boundaries and type coercion for range matching ### Improvements * docs: enable pydoc placeholders for facilitating component reuse through inheritance * docs: change processor natural naming to capital cased with whitespace (e.g. "Generic Resolver") * docs: use processor name placeholders for most usages * docs: hide non-init fields in docs * docs: add examples for generic_resolver * docs: improve example rendering * getter: handle "text/yaml" in content type resolution * getter: remove noisy debug log * vuln: bump aiohttp to at least 3.14.3 in order to fix CVE-2026-69244 * tests: add context handling framework for `test_cases` * tests: add mock_env decorator support for async functions * ci: enforce CHANGELOG.md is updated * ci: enforce PR TODOs are completed * ci: introduce umbrella job for enforcing merge status checks with GitHub * grokker: improve performance by stopping on first matching expression * metrics: cache labeled child metric collectors * metrics: avoid heavy time context manager and hardwire passthrough metric collector methods on wrapper class * metrics: remove `__add__` interface from metrics and use the passthrough instead * metrics: consolidate `measure_time` and `measure_time_async` in single decorator ### Bugfix * chart: fix command handling * ng: fix input timeout to also accept int parameters * ng: fix `http_input` `collect_meta` leading to shared dicts between events * ng: fix error output structure to stay consistent with non-ng * generic_adder: allow `None` as valid input via `add` * grokker: allow fallback matches without named fields * filter: fix lucene range expressions not matching on mixed-type scenarios * filter: treat `inf`/`nan` as string values instead of numeric range boundaries ## 20.0.0 ### Breaking * change `list_comparison` & `network_comparison` processor result to return sub-paths or list names ### Features * calculator: add support for >, <, >=, <=, ==, and != comparisons with arithmetic expressions on both sides * allow `list_comparison` & `network_comparison` to be supplied with explicit list names for results ### Improvements * calculator: add clean expression-stack teardown after parsing or evaluation to safely reuse the shared BNF parser * calculator: refactor the expression grammar into explicit power, multiplicative, additive, and comparison precedence levels * calculator: extend documentation and tests for comparison semantics, chained comparisons, boolean operands, floating-point equality, postfix evaluation order, and parser reuse after failures * add launch configurations for ng and advanced debugging * preprocessor: move configs and partially (ng) logic in dedicated modules * partially use [msgspec.json.Decoder.type](https://github.com/fkie-cad/Logprep/pull/955) for type validation * tests: make tests (automatically) async, refactor component test hierarchy * tests: add testing dependencies (pytest-aiohttp, pytest-timeout) * tests: warm up lazy registry * tests: add ng acceptance and further unit tests * compose: allow setting number of topic partitions using env variables * benchmark: set kafka partitions to 1 for benchmark (configurable) * benchmark: add reference pipelines for ng and non-ng & disable config refresh * benchmark: use kafka-producer over logprep generate * benchmark: make container-runtime configurable (docker/podman) * fix mypy issues * helm: make ng activatable in deployment * performance: avoid dynamically decorating functions with @Metric.measure_time * performance: make component.describe() a cached_property named description * ng: migrate to async * ng: replace iter-based workflow with asynchronous workers & queues * ng: implement graceful worker shutdown exhausting all work queues in topological order * ng: allow finite inputs to be exhausted, leading to graceful application shutdown * ng: rework event type hierarchy, remove states and event backlog * ng: integrate http input server in async loop & use for health check * ng: make error event first citizen (e.g. part of the input interface) * ng: use async client for confluent_kafka input & commit only acknowledged offsets * ng: use async client for confluent_kafka output * ng: remove admin client for confluent_kafka input/output * ng: use async client for opensearch output, *temporarily disabling retries for now* * ng: rework outputs to operate on batches and flush immediately * ng: use `OutputSpec` instead of `list[dict]` to specify desired output targets for extra events * ng: drop "ng_" prefix completely and use a setting in the global registry to activate ng * ng: decouple console logging via queue/listener in separate thread * improve `list_comparison` performance * ng: implement circuit breaker utility to protect downstream systems * ng: enable opensearch output to retry on transport- and item-level failures + circuit breaker * ng: add metrics for opensearch output * improve env var access performance by using a cached snapshot * getter: add debug logs * improve robustness of compose tests ### Bugfix * fix `dissector` not dissecting multiline strings * fix `grokker` dropping matches on duplicate named capture groups * fix `list_comparison` & `network_comparison` to actually work with dotted field notation * fix `list_comparison` & `network_comparison` rule failure handling for multiple lists * fix `list_comparison` & `network_comparison` to raise on http-urls without `LOGPREP_LIST` * fix `list_comparison` & `network_comparison` to raise on ambigous basenames for `list_file_paths` ## 19.4.1 ### Breaking ### Features * add support for `dict` in `generic_adder` ### Improvements * fix Pyparsing deprecation warnings ### Bugfix * fix `add_fields_to` injecting identical objects instead of copies * fix `generic_adder` accumulating state by inserting identical objects in events ## 19.4.0 ### Breaking ### Features * allow `list_comparison` and `network_comparison` to use dynamic values from event to resolve list uris ### Improvements ### Bugfix * make `list_search_base_path` actually optional and correctly overrideable by rule config * add missing documentation for the `network_comparison` processor ## 19.3.0 ### Breaking ### Features * add Lucene range support for integer, floating-point, and lexicographic string values, including quoted ISO-8601 timestamps with timezone offsets. * add `deduplicator` processor that removes duplicates from lists ### Improvements * add test cases for `decoder` and `timestamper` behavior when handling empty messages * remove unused dependencies and move development and documentation-only dependencies to their corresponding optional dependency groups. * lazy-load processor registry components to reduce import overhead when loading the registry module ### Bugfix ## 19.2.0 ### Breaking * chart: change default deployment strategy to RollingUpdate ### Features * allow `list_comparison` and `network_comparison` to extract list values from configurable JSON fields via `content_field` * add basic `FileGetter` content type support for `.txt`, `.json`, and `.yml` files ### Improvements * remove pre-releases * upgrade transitive dependencies * fix #904 by upgrading opensearch-py to 3.2.0 * fix urllib vulnerabilities CVE-2026-44431 & CVE-2026-44432 * make deployment strategy configurable * allow `list_comparison` and `network_comparison` to continue processing with failure tags after getter errors ### Bugfix * fix `decoder` to support special characters in logfmt keys * fix `decoder` to support empty values in logfmt items ## 19.1.0 ### Breaking ### Features * add content-type aware parsing to getters * add support for JSON list sources in `list_comparison` * allow `list_comparison` to load comparison lists from HTTP endpoints returning `application/json` * add content type tracking to refreshable getters ### Improvements * harden GHA workflows against supply chain attacks * centralize trivy cache across branches using a daily workflow * fix vulnerable python dependencies * add .editorconfig for python * add flake.nix and change container building to nix build as well * remove stale rust dependency and code ### Bugfix * properly handle trailing newlines in lucene filters * properly handle lucene filters ending unexpectedly ## 19.0.0 ### Breaking * backslash in filter and many more expressions gets new meaning, breaking filters or rules that use a plain backslash ### Features * make it possible to assign multiple credentials to a single endpoint * support escaping json dot notation in filter queries, processor fields and special processor syntaxes ### Improvements * improve http endpoint security by fully checking basic auth hashes, and doing that in a time constant manner to not expose secrets * improve clusterer performance by removing access via dotted fields where possible * fix several mypy issues ### Bugfix * fix missing examples for processor decoder * fix calculator silently failing on syntax errors * raise TimeParserException if an invalid UNIX timestamp is parsed to prevent timestamper from crashing ## 18.1.0 ### Breaking ### Features * add uv as dependency management, including uv.lock * allow configuration (and auto-creation) of service accounts in helm chart * add new drop_empty flag to allow the `string_splitter` to drop resulting fields that would be empty (e.g. whitespace) * generic_resolver now handles all FieldValue types (including None) ### Improvements * simplify Dockerfile and remove docker build support for `LOGPREP_VERSION` * pytest.param now works with test_cases document generation * fix several mypy issues ### Bugfix * generic_resolver now follows yaml standard and accepts a list instead of relying on the ordering of a dict * generic_resolver now properly handles falsy values in resolve_list and resolve_from_file * decoder errors are handled properly as warnings instead of causing pipeline failures * fix a bug in `dissector` not handling curly braces on the end of a dissect section properly ## 18.0.1 ### Breaking ### Features * headers from incoming http requests can now be copied into events via `copy_headers_to_log` config in http input, `collect_meta` will be deprecated in the future * add new `decoder` processor to decode values from event field, starting with `json`, `base64`, `clf` (see: https://en.wikipedia.org/wiki/Common_Log_Format), `nginx` parser for kubernetes ingress, `syslog_rfc3164`, `syslog_rfc3164_local`, `syslog_rfc5324`, `logfmt`, `cri`, `docker`, `decolorize` (removing color codes in logs) ### Improvements * use follow-imports=silent (instead of skip) to perform more strict type checking * add docs on how to perform memory profiling * make the pipeline example work on MacOS (reduce error queue size) * clean up scheduled jobs and other resources when shutting down components * fix several minor mypy issues and improve static typing ### Bugfix * fix incorrect default-logger lookup by consistently resolving defaults from `DEFAULT_LOG_CONFIG["loggers"]`. * fix a possible race condition in the `geoip_enricher` * fix possible memory leaks in configuration refresh when processors set up scheduled jobs which were not cleaned up ## 18.0.0 ### Breaking * pre detector events now also include host.name if the field value is None ### Features * add support for python 3.14 * allow pre-detector to copy a configurable list of fields from log to detection event * list comparison processor can now also match fields that contain lists in documents * add network comparison processor that can match IPs with networks in CIDR notation ### Improvements * add workflow to partially run & check the compose example * add clarification to `config_refresh_interval` docstring about potential delay under high system load and non-strict timing behavior * mypy checks in the pull request workflow are now applied to the same directories as in the main workflow * update codecov-action from v2 to v5 * add token for codecov workflow * add new test to validate that per-logger log levels correctly override the global log level in LoggerConfig ### Bugfix * fix opensearch output not respecting thread_count config parameter * fix docker-compose and k8s example setups * fix handling of non-string values (e.g. int) as replacement argument for `generic_resolver` * fix documentation for `generic_resolver` rule `append_to_list -> merge_with_target` option * fix grokker using a fixed directory for downloaded patterns, potentially leading to conflicts between processes * fix a bug in the `pre_detector` that could lead to `host.name` of previous events being copied into pre-detections of new events ## 17.0.3 ### Breaking ### Features * implement first prototype of ng logprep runner * ip alerter can now also match fields that contain lists of IPs in documents * make http getters periodically refresh if configured in file path defined by environment variable `LOGPREP_GETTER_CONFIG` * cache http getter results by utilizing the etag header * add per-target (i.e. `localhost:1234/foo`) callbacks to http getters that are called when getters are refreshed with new data * make list comparison processor be refreshable with http getter * make generic adder processor be refreshable with http getter * make generic resolver processor be refreshable with http getter * add option for refreshable getters to return default values if no value could be obtained ### Improvements ### Bugfix * fix error-output not flushing as scheduled ## 17.0.2 ### Breaking ### Features * add `clear_event` field to `add_full_event_to_target_field` ### Improvements * add `acknowledge()` functionality (state change of events and deleting from backlog) * add `event_backlog` to the abstract input interface. * register event in the backlog and return the registered event object. * make `processors` handle Event class based objects * add an EventBacklog class hierarchies * implement an iterator interface to Input connectors * make simple connectors handle Event class based objects * make `opensearch_output` handle Event class based objects * deprecate `s3_output` as it does not fit into new architecture * deprecate `http_output` as it does not fit into new architecture * make confluentkafka_output store Event class based objects * add new class `Pipeline` to ng module * add new class `Sender` to ng module ### Bugfix * fix auto-rule tester getting stuck due to logging ## 17.0.1 ### Breaking ### Features ### Improvements * add ErrorEvent class * add PseudonymEvent Class * add SreEvent class * add LogEvent class * implement abstract Event class to encapsulate event data, processing state, warnings, and errors * integrate dotted field handling methods directly into Event, enabling structured field access and manipulation * support event identity and hashability based on data, allowing usage in sets and as dictionary keys * implement EventState class to manage the lifecycle of log events * integrate a finite state machine to control valid state transitions * add ng packages as namespace in dirs 'unit' and 'logprep' as preparation for new architecture implementation * add abstract EventMetadata class and KafkaInputMetadata class * remove ProcessorResult class in favor of LogEvent class * use LogEvent class in processor base class ### Bugfix * add `@timestamp` field to error documents * fix crash on `TimeParserException` during preprocessing caused by invalid timestamps ## 17.0.0 ### Breaking * removed the deprecated kafka generator. The new generator previously available via the kafka2 CLI has been renamed to kafka. ### Features * add `replacer` processor to replace substrings in fields using a syntax similar to the `dissector` * add custom yaml tag `!include PATH_TO_YAML_FILE` that allows to include other yaml files. * add custom yaml tags `!set_anchor ANCHOR_NAME` and `!load_anchor ANCHOR_NAME` that allow to use anchors across documents inside a file/stream. ### Improvements * ensured that "_test.json" files are not loaded as rules * introduce new logger `Config` * refactor config refresh behavior from `logprep.runner` to `logprep.util.configuration` * refactor config related metrics from `logprep.runner` to `logprep.util.configuration` * added a log message for recovering config refresh mechanic from failing source ### Bugfix * Fixed logging error in _revoke_callback() by adding error handling * Fixed endless loading in logprep test config * prevent the auto rule tester from loading rules directly defined inside the config, since they break the auto rule tester and can't have tests anyways * Fixed typo and broken link in documentation * Fixed assign_callback error in confluentkafka input * Fixed error logging in ` _get_configuration`, which caused the github checks to fail * Resolved `mypy` errors in `BaseProcessorTestCase.` by ensuring `self.object` and `self.patchers` are not `None` before accessing attributes. * Fix domain resolver errors for invalid domains * Fixed deprecation warnings caused by datetime when using Python >= 3.12 * Fixed timestamp and timezone mismatch issue * Fixed a bug where config refresh interval was not reset to original interval after recovering from source related failures (i.e. http timeouts) * Fixed inconsistent generator statistics report during multithreading by making it thread safe ## 16.1.0 ### Deprecations * the generator input config now uses target instead of the deprecated target_path. ### Features * adds new config parameter `event_original_field` to http input which can be used to write the original event in a designated target field * adds a new preprocessor `add_full_event_to_target_field` which adds the full event as an escaped string to a designated target field * added a new confluent kafka generator, can be invoked with the kafka2 command * added --verify option to the generate http command to activate ssl verification and possible set a certificate path ### Improvements * reworked the http generator input class to general input class for http and confluent_kafka * rewrote the input class into seperate classes including batcher, sender, fileloader, input * reworked the generator http controller into a general generator controller * updated documentation for the new generator ### Bugfix * prevent restart timeout for pipelines to rise infinitely ## 16.0.0 ### Breaking * remove `hyperscan_resolver` processor because it is not significantly faster as the `generic_resolver` with enabled cache ### Features * add support for rule files with suffix `.yaml` * add a feature to the preprocessor `log_arrival_time_target_field` to backup the original content on a preexisting target parent field in case of errors during preprocessing ### Improvements * removes `colorama` dependency * reimplemented the rule loading mechanic * removes `rstr` dependency * add mypy to ci * use official python image again and mitigate setuptools related CVE by uninstalling it system wide * refactored code quality pipeline to apply DRY * rewrote pre-detection tests ### Bugfix * fixes a bug with lucene regex and parentheses * fixes a conflict between lucene filter and the Crypto module * fixes error in `_handle_warning_error` that broke up tags into characters if the original tag was not a list * fixes bug in `OAuthClientCredentialsFlow` where the first request session was not closed and overwritten ## 15.1.0 ### Breaking ### Features * add multiarch container builds for AMD64 and ARM64 ### Improvements ### Bugfix ## 15.0.0 ### Breaking * drop support for python 3.10 and add support for python 3.13 * `CriticalInputError` is raised when the input preprocessor values can't be set, this was so far only true for the hmac preprocessor, but is now also applied for all other preprocessors. * fix `delimiter` typo in `StringSplitterRule` configuration * removed the configuration `tld_lists` in `domain_resolver`, `domain_label_extractor` and `pseudonymizer` as the list is now fixed inside the packaged logprep * remove SQL feature from `generic_adder`, fields can only be added from rule config or from file * use a single rule tree instead of a generic and a specific rule tree * replace the `extend_target_list` parameter with `merge_with_target` for improved naming clarity and functionality across `FieldManager` based processors (e.g., `FieldManager`, `Clusterer`, `GenericAdder`). ### Features * configuration of `initContainers` in logprep helm chart is now possible ### Improvements * fix `requester` documentation * replace `BaseException` with `Exception` for custom errors * refactor `generic_resolver` to validate rules on startup instead of application of each rule * regex pattern lists for the `generic_resolver` are pre-compiled * regex matching from lists in the `generic_resolver` is cached * matching in the `generic_resolver` can be case-insensitive * rewrite the helper method `add_field_to` such that it always raises an `FieldExistsWarning` instead of return a bool. * add new helper method `add_fields_to` to directly add multiple fields to one event * refactored some processors to make use of the new helper methods * add `pre-commit` hooks to the repository, install new dev dependency and run `pre-commit install` in the root dir * the default `securityContext`for the pod is now configurable * allow `TimeParser` to get the current time with a specified timezone instead of always using local time and setting the timezone to UTC * remove `tldextract` dependency * remove `urlextract` dependency * fix wrong documentation for `timestamp_differ` * add container signatures to images build in ci pipeline * add sbom to images build in ci pipeline * `FieldManager` supports merging dictionaries ### Bugfix * fix `confluent_kafka.store_offsets` if `last_valid_record` is `None`, can happen if a rebalancing happens before the first message was pulled. * fix pseudonymizer cache metrics not updated * fix incorrect timezones for log arrival time and delta time in input preprocessing * fix `_get_value` in `FilterExpression` so that keys don't match on values * fix `auto_rule_tester` to work with `LOGPREP_BYPASS_RULE_TREE` enabled * fix `opensearch_output` not draining `message_backlog` on shutdown * silence `FieldExists` warning in metrics when `LOGPREP_APPEND_MEASUREMENT_TO_EVENT` is active ## 14.0.0 ### Breaking * remove AutoRuleCorpusTester * removes the option to use synchronous `bulk` or `parallel_bulk` operation in favor of `parallel_bulk` in `opensearch_output` * reimplement error handling by introducing the option to configure an error output * if no error output is configured, failed event will be dropped ### Features * adds health check endpoint to metrics on path `/health` * changes helm chart to use new readiness check * adds `healthcheck_timeout` option to all components to tweak the timeout of healthchecks * adds `desired_cluster_status` option to opensearch output to signal healthy cluster status * initially run health checks on setup for every configured component * make `imagePullPolicy` configurable for helm chart deployments * it is now possible to use Lucene compliant Filter Expressions * make `terminationGracePeriodSeconds` configurable in helm chart values * adds ability to configure error output * adds option `default_op_type` to `opensearch_output` connector to set the default operation for indexing documents (default: index) * adds option `max_chunk_bytes` to `opensearch_output` connector to set the maximum size of the request in bytes (default: 100MB) * adds option `error_backlog_size` to logprep configuration to configure the queue size of the error queue * the opensearch default index is now only used for processed events, errors will be written to the error output, if configured ### Improvements * remove AutoRuleCorpusTester * adds support for rust extension development * adds prebuilt wheels for architectures `x86_64` on `manylinux` and `musllinux` based linux platforms to releases * add manual how to use local images with minikube example setup to documentation * move `Configuration` to top level of documentation * add `CONTRIBUTING` file * sets the default for `flush_timeout` and `send_timeout` in `kafka_output` connector to `0` seconds * changed python base image for logprep to `bitnami/python` in cause of better CVE governance ### Bugfix * ensure `logprep.abc.Component.Config` is immutable and can be applied multiple times * remove lost callback reassign behavior from `kafka_input` connector * remove manual commit option from `kafka_input` connector * pin `mysql-connector-python` to >=9.1.0 to accommodate for CVE-2024-21272 and update `MySQLConnector` to work with the new version ## 13.1.2 ### Bugfix * fixes a bug not increasing but decreasing timeout throttle factor of ThrottlingQueue * handle DecodeError and unexpected Exceptions on requests in `http_input` separately * fixes unbound local error in http input connector ## 13.1.1 ### Improvements * adds ability to bypass the processing of events if there is no pipeline. This is useful for pure connector deployments. * adds experimental feature to bypass the rule tree by setting `LOGPREP_BYPASS_RULE_TREE` environment variable ### Bugfix * fixes a bug in the `http_output` used by the http generator, where the timeout parameter does only set the read_timeout not the write_timeout * fixes a bug in the `http_input` not handling decode errors ## 13.1.0 ### Features * `pre_detector` now normalizes timestamps with configurable parameters timestamp_field, source_format, source_timezone and target_timezone * `pre_detector` now writes tags in failure cases * `ProcessingWarnings` now can write `tags` to the event * add `timeout` parameter to logprep http generator to set the timeout in seconds for requests * add primitive rate limiting to `http_input` connector ### Improvements * switch to `uvloop` as default loop for the used threaded http uvicorn server * switch to `httptools` as default http implementation for the used threaded http uvicorn server ### Bugfix * remove redundant chart features for mounting secrets ## 13.0.1 ### Improvements * a result object was added to processors and pipelines * each processor returns an object including the processor name, generated extra_data, warnings and errors * the pipeline returns an object with the list of all processor result objects * add kubernetes opensiem deployment example * move quickstart setup to compose example ### Bugfix * This release limits the mysql-connector-python dependency to have version less the 9 ## 13.0.0 ### Breaking * This release limits the maximum python version to `3.12.3` because of the issue [#612](https://github.com/fkie-cad/Logprep/issues/612). * Remove `normalizer` processor, as it's functionality was replaced by the `grokker`, `timestamper` and `field_manager` processors * Remove `elasticsearch_output` connector to reduce maintenance effort ### Features * add a helm chart to install logprep in kubernetes based environments ### Improvements * add documentation about behavior of the `timestamper` on `ISO8601` and `UNIX` time parsing * add unit tests for helm chart templates * add helm to github actions runner * add helm chart release to release pipeline ### Bugfix * fixes a bug where it could happen that a config value could be overwritten by a default in a later configuration in a multi source config scenario * fixes a bug in the `field_manager` where extending a non list target leads to a processing failure * fixes a bug in `pseudonymizer` where a missing regex_mapping from an existing config_file causes logprep to crash continuously ## 12.0.0 ### Breaking * `pseudonymizer` change rule config field `pseudonyms` to `mapping` * `clusterer` change rule config field `target` to `source_fields` * `generic_resolver` change rule config field `append_to_list` to `extend_target_list` * `hyperscan_resolver` change rule config field `append_to_list` to `extend_target_list` * `calculator` now adds the error tag `_calculator_missing_field_warning` to the events tag field instead of `_calculator_failure` in case of missing field in events * `domain_label_extractor` now writes `_domain_label_extractor_missing_field_warning` tag to event tags in case of missing fields * `geoip_enricher` now writes `_geoip_enricher_missing_field_warning` tag to event tags in case of missing fields * `grokker` now writes `_grokker_missing_field_warning` tag to event tags instead of `_grokker_failure` in case of missing fields * `requester` now writes `_requester_missing_field_warning` tag to event tags instead of `_requester_failure` in case of missing fields * `timestamp_differ` now writes `_timestamp_differ_missing_field_warning` tag to event tags instead of `_timestamp_differ_failure` in case of missing fields * `timestamper` now writes `_timestamper_missing_field_warning` tag to event tags instead of `_timestamper_failure` in case of missing fields * rename `--thread_count` parameter to `--thread-count` in http generator * removed `--report` parameter and feature from http generator * when using `extend_target_list` in the `field manager`the ordering of the given source fields is now preserved * logprep now exits with a negative exit code if pipeline restart fails 5 times * this was implemented because further restart behavior should be configured on level of a system init service or container orchestrating service like k8s * the `restart_count` parameter is configurable. If you want the old behavior back, you can set this parameter to a negative number * logprep now exits with a exit code of 2 on configuration errors ### Features * add UCL into the quickstart setup * add logprep http output connector * add pseudonymization tools to logprep -> see: `logprep pseudo --help` * add `restart_count` parameter to configuration * add option `mode` to `pseudonymizer` processor and to pseudonymization tools to chose the AES Mode for encryption and decryption * add retry mechanism to opensearch parallel bulk, if opensearch returns 429 `rejected_execution_exception` ### Improvements * remove logger from Components and Factory signatures * align processor architecture to use methods like `write_to_target`, `add_field_to` and `get_dotted_field_value` when reading and writing from and to events * required substantial refactoring of the `hyperscan_resolver`, `generic_resolver` and `template_replacer` * change `pseudonymizer`, `pre_detector`, `selective_extractor` processors and `pipeline` to handle `extra_data` the same way * refactor `clusterer`, `pre_detector` and `pseudonymizer` processors and change `rule_tree` so that the processor do not require `process` override * required substantial refactoring of the `clusterer` * handle missing fields in processors via `_handle_missing_fields` from the field_manager * add `LogprepMPQueueListener` to outsource logging to a separate process * add a single `Queuehandler` to root logger to ensure all logs were handled by `LogprepMPQueueListener` * refactor `http_generator` to use a logprep http output connector * ensure all `cached_properties` are populated during setup time ### Bugfix * make `--username` and `--password` parameters optional in http generator * fixes a bug where `FileNotFoundError` is raised during processing ## 11.3.0 ### Features * add gzip handling to `http_input` connector * adds advanced logging configuration * add configurable log format * add configurable datetime formate in logs * makes `hostname` available in custom log formats * add fine grained log level configuration for every logger instance ### Improvements * rename `logprep.event_generator` module to `logprep.generator` * shorten logger instance names ### Bugfix * fixes exposing OpenSearch/ElasticSearch stacktraces in log when errors happen by making loglevel configurable for loggers `opensearch` and `elasticsearch` * fixes the logprep quickstart profile ## 11.2.1 ### Bugfix * fixes bug, that leads to spawning exporter http server always on localhost ## 11.2.0 ### Features * expose metrics via uvicorn webserver * makes all uvicorn configuration options possible * add security best practices to server configuration * add following metrics to `http_input` connector * `nummer_of_http_requests` * `message_backlog_size` ### Bugfix * fixes a bug in grokker rules, where common field prefixes wasn't possible * fixes bug where missing key in credentials file leads to AttributeError ## 11.1.0 ### Features * new documentation part with security best practices which compiles to `user_manual/security/best_practices.html` * also comes with excel export functionality of given best practices * add basic auth to http_input ### Bugfix * fixes a bug in http connector leading to only first process working * fixes the broken gracefull shutdown behaviour ## 11.0.1 ### Bugfix * fixes a bug where the pipeline index increases on every restart of a failed pipeline * fixes closed log queue issue by run logging in an extra process ## 11.0.0 ### Breaking * configuration of Authentication for getters is now done by new introduced credentials file ### Features * introducing an additional file to define the credentials for every configuration source * retrieve oauth token automatically from different oauth endpoints * retrieve configruation with mTLS authentication * reimplementation of HTTP Input Connector with following Features: * Wildcard based HTTP Request routing * Regex based HTTP Request routing * Improvements in thread-based runtime * Configuration and possibility to add metadata ### Improvements * remove `versioneer` dependency in favor of `setuptools-scm` ### Bugfix * fix version string of release versions * fix version string of container builds for feature branches * fix merge of config versions for multiple configs ## v10.0.4 ### Improvements * refactor logprep build process and requirements management ### Bugfix * fix `generic_adder` not creating new field from type `list` ## v10.0.3 ### Bugfix * fix loading of configuration inside the `AutoRuleCorpusTester` for `logprep test integration` * fix auto rule tester (`test unit`), which was broken after adding support for multiple configuration files and resolving paths in configuration files ## v10.0.2 ### Bugfix * fix versioneer import * fix logprep does not complain about missing PROMETHEUS_MULTIPROC_DIR ## v10.0.1 ### Bugfix * fix entrypoint in `setup.py` that corrupted the install ## v10.0.0 ### Breaking * reimplement the logprep CLI, see `logprep --help` for more information. * remove feature to reload configuration by sending signal `SIGUSR1` * remove feature to validate rules because it is already included in `logprep test config` ### Features * add a `number_of_successful_writes` metric to the s3 connector, which counts how many events were successfully written to s3 * make the s3 connector work with the new `_write_backlog` method introduced by the `confluent_kafka` commit bugfix in v9.0.0 * add option to Opensearch Output Connector to use parallel bulk implementation (default is True) * add feature to logprep to load config from multiple sources (files or uris) * add feature to logprep to print the resulting configruation with `logprep print json|yaml ` in json or yaml * add an event generator that can send records to Kafka using data from a file or from Kafka * add an event generator that can send records to a HTTP endpoint using data from local dataset ### Improvements * a do nothing option do dummy output to ensure dummy does not fill memory * make the s3 connector raise `FatalOutputError` instead of warnings * make the s3 connector blocking by removing threading * revert the change from v9.0.0 to always check the existence of a field for negated key-value based lucene filter expressions * make store_custom in s3, opensearch and elasticsearch connector not call `batch_finished_callback` to prevent data loss that could be caused by partially processed events * remove the `schema_and_rule_checker` module * rewrite Logprep Configuration object see documentation for more details * rewrite Runner * delete MultiProcessingPipeline class to simplify multiprocesing * add FDA to the quickstart setup * bump versions for `fastapi` and `aiohttp` to address CVEs ### Bugfix * make the s3 connector actually use the `max_retries` parameter * fixed a bug which leads to a `FatalOutputError` on handling `CriticalInputError` in pipeline ## v9.0.3 ### Breaking ### Features * make `thread_count`, `queue_size` and `chunk_size` configurable for `parallel_bulk` in opensearch output connector ### Improvements ### Bugfix * fix `parallel_bulk` implementation not delivering messages to opensearch ## v9.0.2 ### Bugfix * remove duplicate pseudonyms in extra outputs of pseudonymizer ## v9.0.1 ### Breaking ### Features ### Improvements * use parallel_bulk api for opensearch output connector ### Bugfix ## v9.0.0 ### Breaking * remove possibility to inject auth credentials via url string, because of the risk leaking credentials in logs - if you want to use basic auth, then you have to set the environment variables * :code:`LOGPREP_CONFIG_AUTH_USERNAME=` * :code:`LOGPREP_CONFIG_AUTH_PASSWORD=` - if you want to use oauth, then you have to set the environment variables * :code:`LOGPREP_CONFIG_AUTH_TOKEN=` * :code:`LOGPREP_CONFIG_AUTH_METHOD=oauth` ### Features ### Improvements * improve error message on empty rule filter * reimplemented `pseudonymizer` processor - rewrote tests till 100% coverage - cleaned up code - reimplemented caching using pythons `lru_cache` - add cache metrics - removed `max_caching_days` config option - add `max_cached_pseudonymized_urls` config option which defaults to 1000 - add lru caching for peudonymizatin of urls * improve loading times for the rule tree by optimizing the rule segmentation and sorting * add support for python 3.12 and remove support for python 3.9 * always check the existence of a field for negated key-value based lucene filter expressions * add kafka exporter to quickstart setup ### Bugfix * fix the rule tree parsing some rules incorrectly, potentially resulting in more matches * fix `confluent_kafka` commit issue after kafka did some rebalancing, fixes also negative offsets ## v8.0.0 ### Breaking * reimplemented metrics so the former metrics configuration won't work anymore * metric content changed and existent grafana dashboards will break * new rule `id` could possibly break configurations if the same rule is used in both rule trees - can be fixed by adding a unique `id` to each rule or delete the possibly redundant rule ### Features * add possibility to convert hex to int in `calculator` processor with new added function `from_hex` * add metrics on rule level * add grafana example dashboards under `examples/exampledata/config/grafana/dashboards` * add new configuration field `id` for all rules to identify rules in metrics and logs - if no `id` is given, the `id` will be generated in a stable way - add verification of rule `id` uniqueness on processor level over both rule trees to ensure metrics are counted correctly on rule level ### Improvements * reimplemented prometheus metrics exporter to provide gauges, histograms and counter metrics * removed shared counter, because it is redundant to the metrics * get exception stack trace by setting environment variable `DEBUG` ### Bugfix ## v7.0.0 ### Breaking * removed metric file target * move kafka config options to `kafka_config` dictionary for `confluent_kafka_input` and `confluent_kafka_output` connectors ### Features * add a preprocessor to enrich by systems env variables * add option to define rules inline in pipeline config under processor configs `generic_rules` or `specific_rules` * add option to `field_manager` to ignore missing source fields to suppress warnings and failure tags * add ignore_missing_source_fields behavior to `calculator`, `concatenator`, `dissector`, `grokker`, `ip_informer`, `selective_extractor` * kafka input connector - implemented manual commit behaviour if `enable.auto.commit: false` - implemented on_commit callback to check for errors during commit - implemented statistics callback to collect metrics from underlying librdkafka library - implemented per partition offset metrics - get logs and handle errors from underlying librdkafka library * kafka output connector - implemented statistics callback to collect metrics from underlying librdkafka library - get logs and handle errors from underlying librdkafka library ### Improvements * `pre_detector` processor now adds the field `creation_timestamp` to pre-detections. It contains the time at which a pre-detection was created by the processor. * add `prometheus` and `grafana` to the quickstart setup to support development * provide confluent kafka test setup to run tests against a real kafka cluster ### Bugfix * fix CVE-2023-37920 Removal of e-Tugra root certificate * fix CVE-2023-43804 `Cookie` HTTP header isn't stripped on cross-origin redirects * fix CVE-2023-37276 aiohttp.web.Application vulnerable to HTTP request smuggling via llhttp HTTP request parser ## v6.8.1 ### Bugfix * Fix writing time measurements into the event after the deleter has deleted the event. The bug only happened when the `metrics.measure_time.append_to_event` configuration was set to `true`. * Fix memory leak by removing the log aggregation capability ## v6.8.0 ### Features * Add option to repeat input documents for the following connectors: `DummyInput`, `JsonInput`, `JsonlInput`. This enables easier debugging by introducing a continues input stream of documents. ### Bugfix * Fix restarting of logprep every time the kafka input connector receives events that aren't valid json documents. Now the documents will be written to the error output. * Fix ProcessCounter to actually print counts periodically and not only once events are processed ## v6.7.0 ### Improvements * Print logprep warnings in the rule corpus tester only in the detailed reports instead of the summary. ### Bugfix * Fix error when writing too large documents into Opensearch/Elasticsearch * Fix dissector pattern that end with a dissect, e.g `system_%{type}` * Handle long-running grok pattern in the `Grokker` by introducing a timeout limit of one second * Fix time handling: If no year is given assume the current year instead of 1900 and convert time zone only once ## v6.6.0 ### Improvements * Replace rule_filter with lucene_filter in predetector output. The old internal logprep rule representation is not present anymore in the predetector output, the name `rule_filter` will stay in place of the `lucene_filter` name. * 'amides' processor now stores confidence values of processed events in the `amides.confidence` field. In case of positive detection results, rule attributions are now inserted in the `amides.attributions` field. ### Bugfix * Fix lucene rule filter representation such that it is aligned with opensearch lucene query syntax * Fix grok pattern `UNIXPATH` by internally converting `[[:alnum:]]` to `\w"` * Fix overwriting of temporary tld-list with empty content ## v6.5.1 ### Bugfix * Fix creation of logprep temp dir * Fix `dry_runner` to support extra outputs of the `selective_extractor` ## v6.5.0 ### Improvements * Make the `PROMETHEUS_MULTIPROC_DIR` environment variable optional, will default to `/tmp/PROMETHEUS_MULTIPROC_DIR` if not given ### Bugfix * All temp files will now be stored inside the systems default temp directory ## v6.4.0 ### Improvements * Bump `requests` to `>=2.31.0` to circumvent `CVE-2023-32681` * Include a lucene representation of the rule filter into the predetector results. The representation is not completely lucene compatible due to non-existing regex functionality. ### Bugfix * Fix error handling of FieldManager if no mapped source field exists in the event. * Fix Grokker such that only the first grok pattern match is applied instead of all matching pattern * Fix Grokker such that nested parentheses in oniguruma pattern are working (3 levels are supported now) * Fix Grokker such that two or more oniguruma can point to the same target. This ensures grok-pattern compatibility with the normalizer and other grok tools ## v6.3.0 ### Features * Extend dissector such that it can trim characters around dissected field with `%{field-( )}` notation. * Extend timestamper such that it can take multiple source_formats. First format that matches will be used, all following formats will be ignored ### Improvements * Extend the `FieldManager` such that it can move/copy multiple source fields into multiple targets inside one rule. ### Bugfix * Fix error handling of missing source fields in grokker * Fix using same output fields in list of grok pattern in grokker ## v6.2.0 ### Features * add `timestamper` processor to extract timestamp functionality from normalizer ### Improvements * removed `arrow` dependency and depending features for performance reasons * switched to `datetime.strftime` syntax in `timestamp_differ`, `s3_output`, `elasticsearch_output` and `opensearch_output` * encapsulate time related functionality in `logprep.util.time.TimeParser` ### Bugfix * Fix missing default grok patterns in packaged logprep version ## v6.1.0 ### Features * Add `amides` processor to extends conventional rule matching by applying machine learning components * Add `grokker` processor to extract grok functionality from normalizer * `Normalizer` writes failure tags if nomalization fails * Add `flush_timeout` to `opensearch` and `elasticsearch` outputs to ensure message delivery within a configurable period * add `kafka_config` option to `confluent_kafka_input` and `confluent_kafka_output` connectors to provide additional config options to `librdkafka` ### Improvements * Harmonize error messages and handling for processors and connectors * Add ability to schedule periodic tasks to all components * Improve performance of pipeline processing by switching form builtin `json` to `msgspec` in pipeline and kafka connectors * Rewrite quickstart setup: * Remove logstash, replace elasticsearch by opensearch and use logprep opensearch connector to stick to reference architecture * Use kafka without zookeeper and switch to bitnami container images ### Bugfix * Fix resetting processor caches in the `auto_rule_corpus_tester` by initializing all processors between test cases. * Fix processing of generic rules after there was an error inside the specific rules. * Remove coordinate fields from results of the geoip enricher if one of them has `None` values ## v6.0.0 ## Breaking * Remove rules deprecations introduced in `v4.0.0` * Changes rule language of `selective_extractor`, `pseudonymizer`, `pre_detector` to support multiple outputs ### Features * Add `string_splitter` processor to split strings of variable length into lists * Add `ip_informer` processor to enrich events with ip information * Allow running the `Pipeline` in python without input/output connectors * Add `auto_rule_corpus_tester` to test a whole rule corpus against defined expected outputs. * Add shorthand for converting datatypes to `dissector` dissect pattern language * Add support for multiple output connectors * Apply processors multiple times until no new rule matches anymore. This enables applying rules on results of previous rules. ### Improvements * Bump `attrs` to `>=22.2.0` and delete redundant `min_len_validator` * Specify the metric labels for connectors (add name, type and direction as labels) * Rename metric names to clarify their meanings (`logprep_pipeline_number_of_warnings` to `logprep_pipeline_sum_of_processor_warnings` and `logprep_pipeline_number_of_errors` to `logprep_pipeline_sum_of_processor_errors`) ### Bugfix * Fixes a bug that breaks templating config and rule files with environment variables if one or more variables are not set in environment * Fixes a bug for `opensearch_output` and `elasticsearch_output` not handling authentication issues * Fix metric `logprep_pipeline_number_of_processed_events` to actually count the processed events per pipeline * Fix a bug for enrichment with environment variables. Variables must have one of the following prefixes now: `LOGPREP_`, `CI_`, `GITHUB_` or `PYTEST_` ### Improvements * reimplements the `selective_extractor` ## v5.0.1 ### Breaking * drop support for python `3.6`, `3.7`, `3.8` * change default prefix behavior on appending to strings of `dissector` ### Features * Add an `http input connector` that spawns a uvicorn server which parses requests content to events. * Add an `file input connector` that reads generic logfiles. * Provide the possibility to consume lists, rules and configuration from files and http endpoints * Add `requester` processor that enriches by making http requests with field values * Add `calculator` processor to calculate with or without field values * Make output subfields of the `geoip_enricher` configurable by introducing the rule config `customize_target_subfields` * Add a `timestamp_differ` processor that can parse two timestamps and calculate their respective time delta. * Add `config_refresh_interval` configuration option to refresh the configuration on a given timedelta * Add option to `dissector` to use a prefix pattern in dissect language for appending to strings and add the default behavior to append to strings without any prefixed separator ### Improvements * Add support for python `3.10` and `3.11` * Add option to submit a template with `list_search_base_path` config parameter in `list_comparison` processor * Add functionality to `geoip_enricher` to download the geoip-database * Add ability to use environment variables in rules and config * Add list access including slicing to dotted field notation for getting values * Add processor boilerplate generator to help adding new processors ### Bugfixes * Fix count of `number_of_processed_events` metric in `input` connector. Will now only count actual events. ## v4.0.0 ### Breaking * Splitting the general `connector` config into `input` and `output` to compose connector config independendly * Removal of Deprecated Feature: HMAC-Options in the connector consumer options have to be under the subkey `preprocessing` of the `input` processor * Removal of Deprecated Feature: `delete` processor was renamed to `deleter` * Rename `writing_output` connector to `jsonl_output` ### Features * Add an opensearch output connector that can be used to write directly into opensearch. * Add an elasticsearch output connector that can be used to write directly into elasticsearch. * Split connector config into seperate config keys `input` and `output` * Add preprocessing capabillities to all input connectors * Add preprocessor for log_arrival_time * Add preprocessor for log_arrival_timedelta * Add metrics to connectors * Add `concatenator` processor that can combine multiple source fields * Add `dissector` processor that tokinizes messages into new or existing fields * Add `key_checker` processor that checks if all dotted fields from a list are present in the event * Add `field_manager` processor that copies or moves fields and merges lists * Add ability to delete source fields to `concatenator`, `datetime_extractor`, `dissector`, `domain_label_extractor`, `domain_resolver`, `geoip_enricher` and `list_comparison` * Add ability to overwrite target field to `datetime_extractor`, `domain_label_extractor`, `domain_resolver`, `geoip_enricher` and `list_comparison` ### Improvements * Validate connector config on class level via attrs classes * Implement a common interface to all connectors * Refactor connector code * Revise the documentation * Add `sphinxcontrib.datatemplates` and `testcase-renderer` to docs * Reimplement `get_dotted_field_value` helper method which should lead to increased performance * Reimplement `dropper` processor code to improve performance ### Deprecations #### Rule Language * `datetime_extractor.datetime_field` is deprecated. Use `datetime_extractor.source_fields` as list instead. * `datetime_extractor.destination_field` is deprecated. Use `datetime_extractor.target_field` instead. * `delete` is deprecated. Use `deleter.delete` instead. * `domain_label_extractor.target_field` is deprecated. Use `domain_label_extractor.source_fields` as list instead. * `domain_label_extractor.output_field` is deprecated. Use `domain_label_extractor.target_field` instead. * `domain_resolver.source_url_or_domain` is deprecated. Use `domain_resolver.source_fields` as list instead. * `domain_resolver.output_field` is deprecated. Use `domain_resolver.target_field` instead. * `drop` is deprecated. Use `dropper.drop` instead. * `drop_full` is deprecated. Use `dropper.drop_full` instead. * `geoip_enricher.source_ip` is deprecated. Use `geoip_enricher.source_fields` as list instead. * `geoip_enricher.output_field` is deprecated. Use `geoip_enricher.target_field` instead. * `label` is deprecated. Use `labeler.label` instead. * `list_comparison.check_field` is deprecated. Use `list_comparison.source_fields` as list instead. * `list_comparison.output_field` is deprecated. Use `list_comparison.target_field` instead. * `pseudonymize` is deprecated. Use `pseudonymizer.pseudonyms` instead. * `url_fields is` deprecated. Use `pseudonymizer.url_fields` instead. ### Bugfixes * Fix resetting of some metric, e.g. `number_of_matches`. ### Breaking ## v3.3.0 ### Features * Normalizer can now write grok failure fields to an event when no grok pattern matches and if `failure_target_field` is specified in the configuration ### Bugfixes * Fix config validation of the preprocessor `version_info_target_field`. ## v3.2.0 ### Features * Add feature to automatically add version information to all events, configured via the `connector > consumer > preprocessing` configuration * Expose logprep and config version in metric targets * Dry-Run accepts now a single json without brackets for input type `json` ### Improvements * Move the config hmac options to the new subkey `preprocessing`, maintain backward compatibility, but mark old version as deprecated. * Make the generic adder write the SQL table to a file and load it from there instead of loading it from the database for every process of the multiprocessing pipeline. Furthermore, only connect to the SQL database on checking if the database table has changed and the file is stale. This reduces the SQL connections. Before, there was permanently one connection per multiprocessing pipeline active and now there is only one connection per Logprep instance active when accessing the database. ### Bugfixes * Fix SelectiveExtractor output. The internal extracted list wasn't cleared between each event, leading to duplication in the output of the processor. Now the events are cleared such that only the result of the current event is returned. ## v3.1.0 ### Features * Add metric for mean processing time per event for the full pipeline, in addition to per processor ### Bugfixes * Fix performance of the metrics tracking. Due to a store metrics statement at the wrong position the logprep performance was dramatically decreased when tracking metrics was activated. * Fix Auto Rule Tester which tried to access processor stats that do not exist anymore. ## v3.0.0 ### Features * Add ability to add fields from SQL database via GenericAdder * Prometheus Exporter now exports also processor specific metrics * Add `--version` cli argument to print the current logprep version, as well as the configuration version if found ### Improvements * Automatically release logprep on pypi * Configure abstract dependencies for pypi releases * Refactor domain resolver * Refactor `processor_stats` to `metrics`. Metrics are now collected in separate dataclasses ### Bugfixes * Fix processor initialization in auto rule tester * Fix generation of RST-Docs ### Breaking * Metrics refactoring: * The json output format of the previously known status_logger has changed * The configuration key word is now `metrics` instead of `status_logger` * The configuration for the time measurement is now part of the metrics configuration * The metrics tracking still includes values about how many warnings and errors happened, but not of what type. For that the regular logprep logging should be consolidated. ## v2.0.1 ### Bugfixes * Clear matching rules before processing in clusterer * Add missing sphinxcontrib-mermaid in tox.ini ## v2.0.0 ### Features * Add generic processor interface `logprep.abc.processor.Processor` * Add `delete` processor to be used with rules. * Delete `donothing` processor * Add `attrs` based `Config` classes for each processor * Add validation of processor config in config class * Make all processors using python `__slots__` * Add `ProcessorRegistry` to register all processors * Remove plugins feature * Add `ProcessorConfiguration` as an adapter to create configuration for processors * Remove all specific processor factories in favor of `logprep.processor.processor_factory.ProcessorFactory` * Rewrite `ProcessorFactory` * Automate processor configuration documentation * generalize config parameter for using tld lists to `tld_lists` for `domain_resolver`, `domain_label_extractor`, `pseudonymizer` * refactor `domain_resolver` to make code cleaner and increase test coverage ### Bugfixes * remove `ujson` dependency because of CVE