FTDC stands for Full-Time Diagnostic Data Capture. FTDC is used to capture data about the mongod and mongos processes and the system that a mongo process is running on.
From a top down view, An
FTDCController
object lives as a decoration on the service context. It is registered
here. The
FTDCController is initialized from the mongod and mongos main functions, which call
startMongoDFTDC
and
startMongoSFTDC
respectively. The FTDC controller owns a collection of collector objects that gather system and
process information for the controller. These sets of collector objects are stored in a
FTDCCollectorCollection
object, allowing all the data to be collected through one call to collect on the
FTDCCollectorCollection. The controller owns several such collections. The two primary ones are
_periodicCollectors
that collects data at a specified time interval, and
_rotateCollectors
that collects one set of data every time a file is created. There is also a
_periodicMetadataCollectors
collection for slowly-changing metadata that is sampled periodically and delta encoded (see
Metadata Delta encoding).
At specified time intervals, the FTDC Controller calls collect on the _periodicCollectors
collection. To collect the data, most collectors run existing commands via
CommandHelpers::runCommandDirectly
which is abstracted into FTDCSimpleInternalCommandCollector. Some process data is gathered via
custom implementations of FTDCCollectorInterface. After gathering the data, the FTDC Controller
writes
the data out as described below using an
FTDCFileManager.
Most collectors are of the class
FTDCSimpleInternalCommandCollector.
The
FTDCServerStatusCommandCollector.
has a very specific query that it needs to run to get the server information. For example, it does
not collect sharding information from serverStatus because sharding could trigger issues with
compression efficiency of FTDC. The system stats are collected by LinuxSystemMetricsCollector via
/proc on Linux and WindowsSystemMetricsCollector via Windows perf counters on Windows.
The
FTDCFileManager
is responsible for rotating files, managing disk space, and recovering files after a crash. When it
gets a hold of the data to be written to disk, it initially writes the data to an
FTDCCompressor
object. The
FTDCCompressor
compresses the data object and appends it to a chunk to be written out. If the chunk is full or the
BSON Schema has changed, the compressor indicates that the chunk needs to be written out to disk
immediately. Otherwise, the compressor indicates that it can store more data.
If the compressor can store more data there are two possibilities. If the data in the compressor has
reached a certain threshold, the
FTDCFileWriter
will decide to write out the unwritten section of data to an interim file. The goal of the interim
file is to flush data out in smaller intervals between archives to ensure the most up-to-date data
is recorded before a crash. The interim file also provides a chance for the regular files to have
maximum sized and compressed chunks. If not, nothing is written to outside of the compressor. If the
compressor indicates that the chunk needs to be written out immediately, data is flushed out of the
compressor into the archive file. The interim file is erased and written over, and the compressor
resets to take in new values.
The
FTDCFileManager
also decides when to rotate the archive files. When the file gets too large, the manager deletes the
reference to the old file and starts writing immediately to the new file by calling
rotate.
FTDC writes two types of files in diagnostic.data.
Archive files: metrics.%Y-%m-%dT%H-%M-%SZ-CCCCC where
%Y-%m-%dT%H-%M-%SZ-strftimeformat string to format a UTC date time string. Ex:2024-03-08T22-58-41ZCCCCC- is a five digit uniquifier in case multiple files are opened in one second. It is always00000except in unit tests.
It is an append only file which can be read with bsondump. It is composed of several types of bson
documents. See Archive Format. FTDC creates new archive files on server
restart and when the file grows larger then the size cap.
Interm file: metrics.interim
There is always one file and it always has just one document of the metric type. It represents the most recent uncompressed metrics sample. The file is constantly overwritten in contrast to the archive file.
Assumptions:
- All numbers are encoded little endian
- Everything is a BSON document or BSON field unless noted
Using a pseudo EBNF format, the FTDC archive format is as follows:
FTDC = ftdc_doc*
ftdc_doc = metadata
| metric
| metadata_delta
metadata =
_id : DateTime
type: 0
doc : collectors_doc
collectors_doc = // metrics reference doc
start: DateTime
collect_doc+
end: DateTime
collect_doc =
string : { // string is an arbitrary field name here - the collector name
start: DateTime
bson_elements+
end: DateTime
}
metric =
_id : DateTime
type: 1
data : BinData(0) // see metrics_chunk
metrics_chunk = // Not a BSON document, raw bytes
uncompressed_size : uint32_t
compressed_chunk: uint8_t[] // zlib compressed
compressed_chunk = // Not a BSON document, raw bytes
reference_document uint8_t[] // a BSON Document - see collectors_doc
metric_count uint32_t
sample_count uint32_t
compressed_metrics_array uint8_t[] // See compressed chunk format below
metadata_delta =
_id : DateTime
type: 2
count: int32
doc : collectors_doc | collectors_delta_doc
collectors_delta_doc = // collectors_doc excluding unchanged collector sub-documents and unchanged fields
start: DateTime
collect_delta_doc+
end: DateTime
collect_delta_doc = // only present for collectors whose sub-document changed
string : { // the collector name
start: DateTime
changed_bson_elements+
end: DateTime
}Compressed Chunk is a delta, run length encoded, and varint compressed array of numbers.
The FTDC compressors extracts a series of uint64_t numbers from a subset of BSON fields from a
series of BSON documents. All documents in a compressed chunk have the same number of numbers. If
the count changes, the current chunk is closed and a new one is open. As a result of this behavior,
FTDC compression is poor if the number of numerical fields changes. FTDC compression only extracts
numeric fields from certain BSON types. These types are double, int64, int32, boolean,
datetime, decimal128, and timestamp. It is a lossy compression for some types as numbers are
cast to uint64_t.
FTDC extracts these metrics and stores them in a column oriented format. These numbers are then processed in the following steps
- Extraction
- Delta encode
- Run length encoding of zeros
- Varint compression
- ZLIB Compression
For a given BSON document, all numeric fields are extracted and converted to uint64_t. Timestamps
are extracted as two uint64_t, seconds followed by increment. For doubles, NaN is encoded as
zero, doubles less than MIN_INT64 are stored as MIN_INT64, and doubles greater than MAX_INT64
are stored as MAX_INT64.
Often, values do not change between samples (for instance the process id of a process never changes). To reduce the space these values take, we subtract the value of previous sample from the current sample. The first sample after the reference document uses the values from the reference document as its baseline.
A sequence of zeros is compressed to a pair of numbers [0, x] where x is non-zero positive
integer that indicates the number of additional zeros in a sequence. For instance, an array of
zeros [0, 0, 0, 0] is transformed to [0, 3].
Varint encoding is a way to reduce the number of bytes needed to represent an integer. For more information, see VarInt reference. The FTDC reference implementation uses the varint implementation from S2.
The entire block is compressed with the zlib compression algorithm.
For instance, for the following sequence of documents:
Example:
{"a": 1, "x" : 2, "s" : "t"}
{"a": 2, "x" : 2, "s" : "t"}
{"a": 3, "x" : 2, "s" : "t"}
{"a": 4, "x" : 2, "s" : "t"}The first document {"a": 1, "x": 2, "s": "t"} is stored as the reference document. FTDC then
builds an array of [2, 3, 4, 2, 2, 2] to represent the a field followed by the x field.
Next, FTDC computes the delta for each sample in the chunk from the previous chunk. Nothing changes
in the reference document but array is transformed to [1, 1, 1, 0, 0, 0].
Next the array is encoded with run length encoding for zeros [1, 1, 1, 0, 2]. Notice the length of
the array is shorter.
Next, these numbers are written to a block of memory with VarInt encoding. If these numbers are written as is, it would be 8 bytes * 5 numbers for a total of 40 bytes. But by using varint, a smaller number of bytes can be used (5 in this example).
Finally, the chunk is compressed with zlib compression.
The metadata delta files are plain bson but use a simple diff algorithm to de-duplicate data. This format was chosen because it tracks things like server parameters which change infrequently and to preserve string values.
The deltas are computed by comparing the previous periodic metadata sample and the current sample.
The sample is expected to have two levels: top-level start/end timestamps plus one sub-document
per collector, where each sub-document has its own start/end timestamps and fields. The
algorithm compares both levels but is not fully recursive - it does not descend below a collector's
sub-document. If the sequence of field names changes at either level, delta tracking is reset and
the full sample is emitted. The counter returned here is the value written to the count field of
the metadata_delta document.
Python like pseudo code:
previous_doc = None
counter = 0
def delta_doc(doc):
# Reset delta tracking if the top-level field names changed.
if previous_doc is None or not same_field_names(doc, previous_doc):
previous_doc = doc
counter = 0
return (counter, doc)
out_doc = {}
has_changes = False
for name in doc.field_names(): # top level
if name == "start" or name == "end":
out_doc[name] = doc[name] # top-level start/end always copied
continue
sub = doc[name] # collector sub-document
prev_sub = previous_doc[name]
# Reset if this collector's field names changed.
if not same_field_names(sub, prev_sub):
previous_doc = doc
counter = 0
return (counter, doc)
sub_out = {}
sub_changed = False
for field in sub.field_names(): # second level
if field == "start" or field == "end":
sub_out[field] = sub[field] # collected, but only kept if the sub-doc changed
elif sub[field] != prev_sub[field]:
sub_out[field] = sub[field]
sub_changed = True
# Include the collector sub-document only if one of its fields changed.
if sub_changed:
out_doc[name] = sub_out
has_changes = True
if not has_changes:
return None # nothing is emitted for this sample
previous_doc = doc # the baseline is the previous sample
counter += 1
return (counter, out_doc)To reconstruct a periodic metadata sample at a point in time, you can replay all the previous deltas. Treating each document as a dictionary/map, merge each delta into the snapshot, descending one level into the collector sub-documents so that only the changed fields are updated. Note that FTDC does not currently implement this algorithm since it does not need to reconstruct point-in-time snapshots.
Python like pseudo code:
def reconstruct(docs[]):
out_doc = {}
for doc in docs:
for name in doc.field_names():
if name == "start" or name == "end":
out_doc[name] = doc[name]
else:
# merge the collector sub-document field by field
out_doc.setdefault(name, {})
for field in doc[name].field_names():
out_doc[name][field] = doc[name][field]
return out_doc