Internal File Registry Service - This service acts as a registry for the internal location and repre
100K+
This service acts as a registry for the internal location and representation of files.
This service provides functionality to administer files stored in an S3-compatible object storage. All file-related metadata is stored in an internal mongodb database, owned and controlled by this service. It exposes no REST API endpoints and communicates with other services via events.
This event signals that there is a file to register in the database. The file-related metadata from this event gets saved in the database and the file is moved from the incoming staging bucket to the permanent storage.
This event signals that there is a file that needs to be staged for download. The file is then copied from the permanent storage to the outbox for the actual download.
This event is published after a file was registered in the database. It contains all the file-related metadata that was provided by the files_to_register event.
This event is published after a file was successfully staged to the outbox.
We recommend using the provided Docker container.
A pre-built version is available at docker hub:
docker pull ghga/internal-file-registry-service:7.4.0
Or you can build the container yourself from the ./Dockerfile:
# Execute in the repo's root dir:
docker build -t ghga/internal-file-registry-service:7.4.0 .
For production-ready deployment, we recommend using Kubernetes, however, for simple use cases, you could execute the service using docker on a single server:
# The entrypoint is preconfigured:
docker run -p 8080:8080 ghga/internal-file-registry-service:7.4.0 --help
If you prefer not to use containers, you may install the service from source:
# Execute in the repo's root dir:
pip install .
# To run the service:
ifrs --help
The service requires the following configuration parameters:
enable_opentelemetry (boolean): If set to true, this will run necessary setup code.If set to false, no setup code is run, which leaves tracing disabled. Default: false.
otel_trace_sampling_rate (number): Determines which proportion of spans should be sampled. A value of 1.0 means all and is equivalent to the previous behaviour. Setting this to 0 will result in no spans being sampled, but this does not automatically set enable_opentelemetry to False. Minimum: 0. Maximum: 1. Default: 1.0.
log_level (string): The minimum log level to capture. Must be one of: "CRITICAL", "ERROR", "WARNING", "INFO", "DEBUG", or "TRACE". Default: "INFO".
service_instance_id (string, required): A string that uniquely identifies this instance across all instances of this service. A globally unique Kafka client ID will be created by concatenating the service_name and the service_instance_id.
Examples:
"germany-bw-instance-001"
log_format: If set, will replace JSON formatting with the specified string format. If not set, has no effect. In addition to the standard attributes, the following can also be specified: timestamp, service, instance, level, correlation_id, and details. Default: null.
Examples:
"%(timestamp)s - %(service)s - %(level)s - %(message)s"
"%(asctime)s - Severity: %(levelno)s - %(msg)s"
log_traceback (boolean): Whether to include exception tracebacks in log messages. Default: true.
object_storages (object, required): Can contain additional properties.
file_upload_topic (string, required): Topic containing published FileUpload outbox events.
Examples:
"file-uploads"
"file-upload-topic"
file_internally_registered_topic (string, required): Name of the topic used for events indicating that a file has been registered for download.
Examples:
"file-registrations"
"file-registrations-internal"
file_internally_registered_type (string, required): The type used for event indicating that that a file has been registered for download.
Examples:
"file_internally_registered"
file_deleted_topic (string, required): Name of the topic used for events indicating that a file has been deleted.
Examples:
"file-deletions"
file_deleted_type (string, required): The type used for events indicating that a file has been deleted.
Examples:
"file_deleted"
file_staged_topic (string, required): Name of the topic used for events indicating that a new file has been internally registered.
Examples:
"file-stagings"
file_staged_type (string, required): The type used for events indicating that a new file has been internally registered.
Examples:
"file_staged_for_download"
file_deletion_request_topic (string, required): The name of the topic to receive events informing about files to delete.
Examples:
"file-deletion-requests"
file_deletion_request_type (string, required): The type used for events indicating that a request to delete a file has been received.
Examples:
"file_deletion_requested"
files_to_stage_topic (string, required): Name of the topic used for events indicating that a download was requested for a file that is not yet available in the outbox.
Examples:
"file-staging-requests"
files_to_stage_type (string, required): The type used for non-staged file request events.
Examples:
"file_staging_requested"
kafka_servers (array, required): A list of connection strings to connect to Kafka bootstrap servers.
Examples:
[
"localhost:9092"
]
kafka_security_protocol (string): Protocol used to communicate with brokers. Valid values are: PLAINTEXT, SSL. Must be one of: "PLAINTEXT" or "SSL". Default: "PLAINTEXT".
kafka_ssl_cafile (string): Certificate Authority file path containing certificates used to sign broker certificates. If a CA is not specified, the default system CA will be used if found by OpenSSL. Default: "".
kafka_ssl_certfile (string): Optional filename of client certificate, as well as any CA certificates needed to establish the certificate's authenticity. Default: "".
kafka_ssl_keyfile (string): Optional filename containing the client private key. Default: "".
kafka_ssl_password (string, format: password, write-only): Optional password to be used for the client private key. Default: "".
generate_correlation_id (boolean): A flag, which, if False, will result in an error when trying to publish an event without a valid correlation ID set for the context. If True, a new correlation ID will be generated and used in the event header. Default: true.
Examples:
true
false
kafka_max_message_size (integer): The largest message size that can be transmitted, in bytes, before compression. Only services that have a need to send/receive larger messages should set this. When used alongside compression, this value can be set to something greater than the broker's message.max.bytes field, which effectively concerns the compressed message size. Exclusive minimum: 0. Default: 1048576.
Examples:
1048576
16777216
kafka_compression_type: The compression type used for messages. Valid values are: None, gzip, snappy, lz4, and zstd. If None, no compression is applied. This setting is only relevant for the producer and has no effect on the consumer. If set to a value, the producer will compress messages before sending them to the Kafka broker. If unsure, zstd provides a good balance between speed and compression ratio. Default: null.
Examples:
null
"gzip"
"snappy"
"lz4"
"zstd"
kafka_max_retries (integer): The maximum number of times to immediately retry consuming an event upon failure. Works independently of the dead letter queue. Minimum: 0. Default: 0.
Examples:
0
1
2
3
5
kafka_enable_dlq (boolean): A flag to toggle the dead letter queue. If set to False, the service will crash upon exhausting retries instead of publishing events to the DLQ. If set to True, the service will publish events to the DLQ topic after exhausting all retries. Default: false.
Examples:
true
false
kafka_dlq_topic (string): The name of the topic used to resolve error-causing events. Default: "dlq".
Examples:
"dlq"
kafka_retry_backoff (integer): The number of seconds to wait before retrying a failed event. The backoff time is doubled for each retry attempt. Minimum: 0. Default: 0.
Examples:
0
1
2
3
5
mongo_dsn (string, format: multi-host-uri, required): MongoDB connection string. Might include credentials. For more information see: https://naiveskill.com/mongodb-connection-string/. Length must be at least 1.
Examples:
"mongodb://localhost:27017"
db_name (string, required): Name of the database located on the MongoDB server.
Examples:
"my-database"
mongo_timeout: Timeout in seconds for API calls to MongoDB. The timeout applies to all steps needed to complete the operation, including server selection, connection checkout, serialization, and server-side execution. When the timeout expires, PyMongo raises a timeout exception. If set to None, the operation will not time out (default MongoDB behavior). Default: null.
Examples:
300
600
null
db_version_collection (string, required): The name of the collection containing DB version information for this service.
Examples:
"ifrsDbVersions"
migration_wait_sec (integer, required): The number of seconds to wait before checking the DB version again.
Examples:
5
30
180
migration_max_wait_sec: The maximum number of seconds to wait for migrations to complete before raising an error. Default: null.
Examples:
null
300
600
3600
S3Config (object): S3-specific config params.
Inherit your config class from this class if you need to talk
to an S3 service in the backend. Cannot contain additional properties.
s3_endpoint_url (string, required): URL to the S3 API.
Examples:
"http://localhost:4566"
s3_access_key_id (string, required): Part of credentials for login into the S3 service. See: https://boto3.amazonaws.com/v1/documentation/api/latest/guide/credentials.html.
Examples:
"my-access-key-id"
s3_secret_access_key (string, format: password, required and write-only): Part of credentials for login into the S3 service. See: https://boto3.amazonaws.com/v1/documentation/api/latest/guide/credentials.html.
Examples:
"my-secret-access-key"
s3_session_token: Part of credentials for login into the S3 service. See: https://boto3.amazonaws.com/v1/documentation/api/latest/guide/credentials.html. Default: null.
Examples:
"my-session-token"
aws_config_ini: Path to a config file for specifying more advanced S3 parameters. This should follow the format described here: https://boto3.amazonaws.com/v1/documentation/api/latest/guide/configuration.html#using-a-configuration-file. Default: null.
Examples:
"~/.aws/config"
S3ObjectStorageNodeConfig (object): Configuration for one specific object storage node and one bucket in it.
The bucket is the main bucket that the service is responsible for. Cannot contain additional properties.
credentials (required): Refer to #/$defs/S3Config.
A template YAML for configuring the service can be found at
./example-config.yaml.
Please adapt it, rename it to .ifrs.yaml, and place it in one of the following locations:
./.ifrs.yaml)~/.ifrs.yaml)The config yaml will be automatically parsed by the service.
Important: If you are using containers, the locations refer to paths within the container.
All parameters mentioned in the ./example-config.yaml
could also be set using environment variables or file secrets.
For naming the environment variables, just prefix the parameter name with ifrs_,
e.g. for the host set an environment variable named ifrs_host
(you may use both upper or lower cases, however, it is standard to define all env
variables in upper cases).
To use file secrets, please refer to the corresponding section of the pydantic documentation.
This is a Python-based service following the Triple Hexagonal Architecture pattern. It uses protocol/provider pairs and dependency injection mechanisms provided by the hexkit library.
For setting up the development environment, we rely on the devcontainer feature of VS Code in combination with Docker Compose.
To use it, you have to have Docker Compose as well as VS Code with its "Remote - Containers"
extension (ms-vscode-remote.remote-containers) installed.
Then open this repository in VS Code and run the command
Remote-Containers: Reopen in Container from the VS Code "Command Palette".
This will give you a full-fledged, pre-configured development environment including:
Moreover, inside the devcontainer, a command dev_install is available for convenience.
It installs the service with all development dependencies, and it installs pre-commit.
The installation is performed automatically when you build the devcontainer. However,
if you update dependencies in the ./pyproject.toml or the
./requirements-dev.txt, please run it again.
This repository is free to use and modify according to the Apache 2.0 License.
This README file is auto-generated, please see readme_generation.md
for details.
Content type
Image
Digest
sha256:af23b62b6…
Size
76.1 MB
Last updated
11 days ago
docker pull ghga/ifrs:15.3.1-rc.5Pulls:
7,019
Last week