Skip to content

Tags: txn2/mcp-datahub

Tags

v1.15.0

Toggle v1.15.0's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
feat(client): glossary node create and hierarchy enumeration (#199) (#…

…200)

The client could operate on individual glossary terms but never answer
"what is the shape of the glossary?" — no root listing, no children of a
node, no parent chain, and no way to create a node at all. A glossary is
a tree by design, so a consumer could only ever see a flat slice of it.

Adds:

  CreateGlossaryNode(ctx, name, definition, parentNode)
  GetRootGlossaryNodes(ctx, start, count)
  GetRootGlossaryTerms(ctx, start, count)
  GetGlossaryNodeChildren(ctx, nodeURN, start, count)
  GetGlossaryParentChain(ctx, urn)

plus types.GlossaryNode and types.GlossaryChildren. No new MCP tools;
existing term operations are unchanged.

How children are enumerated was the open question in the issue, since
RelationshipsInput.types is a free-form string list and the schema names
no relationship for glossary parentage. Settled against a live DataHub
v1.6.0 rather than guessed: children are the INCOMING side of the
IsPartOf relationship on the parent node. Verified on a real tree — the
edge returns both nodes and terms, pages on start/count, and its total
matches the node's own childrenCount. This matches how DataHub's own UI
fetches node children (datahub-web-react/src/graphql/glossaryNode.graphql).

Two behaviours only the live instance revealed:

- Children lag writes. They come from the graph index, which DataHub
  populates asynchronously, so a just-created child is not visible at
  once. Documented on the method; the integration test polls.
- An unknown node returns an empty stub rather than an error, which is
  indistinguishable from a childless node, so the client selects exists
  and returns ErrNotFound. It is read through a pointer, so a DataHub
  version that omits the field is not misread as absent.

GetGlossaryParentChain reads parentNodes on the entity itself and is
immediately consistent. It returns the chain direct-parent first, matching
DataHub's order, and fills each node's ParentNode from the next link so a
caller can rebuild the branch without another round trip. Non-glossary
URNs are rejected with ErrInvalidURN rather than silently returning an
empty chain.

No change to the support floor: createGlossaryNode, getRootGlossaryNodes,
getRootGlossaryTerms, parentNodes, and childrenCount are all present in
entity.graphql at the v1.3.0 tag, the documented minimum.

Also repairs write_integration_test.go, which no longer compiled: getAspect,
readGlobalTags, readGlossaryTerms, and readInstitutionalMemory had each
gained an entityType parameter that the integration tests were never updated
for. The full integration suite now builds and passes against v1.6.0.

Verified: make verify clean (lint 0 issues, coverage 93.0%, client 95.1%);
make test-integration green against a live DataHub v1.6.0, including the new
TestIntegrationGlossaryHierarchy, which builds a glossary tree, enumerates it
every way the issue asks for, and deletes it.

v1.13.0

Toggle v1.13.0's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
fix(tools): resolve description text from value or description (#194) (

…#195)

datahub_update with what=description or what=column_description took the new
text from the value parameter only, while the same input schema exposes a
description parameter for what=query, incident, and structured_property. Text
passed as description left value empty, so the handler wrote an empty string and
answered {"action":"updated"}.

The empty write is destructive: editableFieldInfo.Description is omitempty, so
the field entry is rewritten without a description and whatever the column
carried is gone.

Both description handlers now resolve the text through resolveDescription:
value is authoritative, description is accepted when value is empty, a genuine
disagreement between the two is rejected, and an update carrying no text at all
is refused rather than written. The value and description parameter descriptions
state which what values use which field.

v1.12.0

Toggle v1.12.0's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
ci: roll up dependabot action bumps (#175, #176, #177, #178) (#183)

* ci: roll up dependabot action bumps (#175, #176, #177, #178)

- github/codeql-action/autobuild 4.36.2 -> 4.37.0 (#175)
- docker/setup-buildx-action 4.1.0 -> 4.2.0 (#176)
- github/codeql-action/upload-sarif 4.36.2 -> 4.37.0 (#177)
- github/codeql-action/analyze 4.36.2 -> 4.37.0 (#178)

* ci: bump codeql-action/init to 4.37.0 to match autobuild/analyze

All codeql-action sub-actions (init, autobuild, analyze) must run the same
version; the init line was left at 4.36.2, causing autobuild to fail with
'Loaded a configuration file for version 4.36.2, but running version 4.37.0'.

v1.11.0

Toggle v1.11.0's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
feat(client): tag descriptions and custom properties (#173) (#174)

* fix(client): correct v3 UPSERT and v1 aspect-read envelope

Two defects surfaced by live testing against DataHub v1.6.0:

- postAspectV3 wrote with default CREATE semantics, so every OpenAPI v3
  read-modify-write (UpdateDescription, AddTag, etc.) returned HTTP 400
  ('Cannot perform CREATE since the aspect already exists') whenever the
  target aspect already existed. Append createIfNotExists=false to force
  UPSERT, which both creates and updates.

- getAspect parsed only the v3 {"value":...} envelope. The legacy v1
  Rest.li GET returns {"aspect":{"<FQCN>":...}}, so v1 read-modify-write
  started from an empty aspect and silently dropped existing tags, terms,
  links, and required fields on write-back. Parse the Rest.li envelope
  (tolerating a normalized flat body), covered by a real-envelope test.

Validated on DataHub v1.6.0 across both v1 and v3 paths.

* feat(client): tag descriptions and custom properties (#173)

Add the two capabilities the downstream curation flow needs:

- UpdateDescription now accepts tag URNs, routing them through the GraphQL
  updateDescription mutation (DataHub's UpdateDescriptionResolver handles
  TAG -> tagProperties.description). Previously tags returned
  ErrUnsupportedEntityType.

- New SetCustomProperties / RemoveCustomProperties client methods write the
  legacy customProperties map via REST read-modify-write, preserving all
  other aspect fields. Aspect names verified against upstream PDL; tag is
  excluded because tagProperties has no customProperties field.

Exposed through the datahub_update tool as what=custom_properties
(set/remove), keeping the tool count unchanged.

Validated end-to-end on DataHub v1.6.0 (v1 and v3 paths).

v1.10.2

Toggle v1.10.2's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
fix(client): read tags/terms for GraphQL-only entity types (#168) (#169)

domain, glossaryTerm, and glossaryNode support writing globalTags and
glossaryTerms via the addTag/addTerm GraphQL mutations, but their
associations could not be read back: these types do not expose typed
tags/glossaryTerms GraphQL fields, and the REST aspect API does not
register those aspects for them. GetEntity therefore returned empty
tags/terms, breaking post-write verification and rollback before-images
for downstream consumers.

Populate Tags/GlossaryTerms (by URN) for these three types by reading
the globalTags/glossaryTerms raw aspects through the experimental
aspects API. The read runs as a separate GetEntityAspectsQuery and
degrades gracefully: any failure (older DataHub, unauthorized token)
leaves the associations empty and logs at debug level rather than
failing the whole GetEntity call, matching the existing graceful
fallback pattern used by GetIncidents/GetStructuredProperties.

Note: raw aspect payloads carry only URNs, so Name/Description are not
populated for these types.

v1.10.1

Toggle v1.10.1's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
fix(client): return full document projection from GetRelatedDocuments (

…#166) (#167)

GetRelatedDocumentsQuery requested a hand-written thin subset (urn,
subType, title, contents, status) in each per-entity ... on Document
block, omitting settings{showInGlobalContext} and relatedAssets even
though the relatedDocumentsResult parse struct and mapping already
support them. As a result every related document parsed with
Settings == nil, so a visibility-gated consumer treating absent
settings as visible (DataHub's documented default) would surface a
document a steward explicitly hid via showInGlobalContext=false. The
same document fetched via GetDocument/SearchDocuments carried the flag
and was correctly suppressed, so documents leaked through this path only.

Reuse the shared documentSelectionFields fragment (introduced in #165
for GetDocument/SearchDocuments) in each relatedDocuments block, so
related documents carry the identical projection: settings,
relatedAssets, status, subType, owners, tags, glossary terms, domain.

Add a httptest-mock test asserting a related document returned with
showInGlobalContext=false and a relatedAssets entry parses with
Settings.ShowInGlobalContext == false and the related-asset URN set.

v1.10.0

Toggle v1.10.0's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
feat(client): add SearchDocuments for document discovery (#164) (#165)

Add Client.SearchDocuments(ctx, query, opts...) to discover context
documents by relevance without a known URN, scoped to the DOCUMENT
entity type via searchAcrossEntities. A "*" query lists all documents;
a text query ranks by relevance. Each result carries the same metadata
as GetDocument (URN, title, sub-type, related-asset URNs,
showInGlobalContext, ownership, tags, glossary terms, domain).

This is a library method, not a new MCP tool: it unblocks
txn2/mcp-data-platform#692, which can add a DataHub-documents source to
its unified search so context-document knowledge becomes discoverable.

- Extract shared documentSelectionFields GraphQL fragment reused by
  GetDocumentQuery and SearchDocumentsQuery so single-document reads and
  document search return identical metadata from one source of truth.
- Extract Client.buildBaseSearchInput helper shared by
  doSearchAcrossEntities and SearchDocuments.
- Add EntityTypeDocument constant.
- Register SearchDocumentsQuery in schema validation.
- Cover with GraphQL httptest mock tests (request scoping, result
  fields incl. governance metadata, wildcard listing, limit capping,
  error handling).

v1.9.0

Toggle v1.9.0's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
feat(dataproducts): return constituent datasets from get_data_product (

…#156)

datahub_get_data_product promised "constituent datasets" and "member
datasets" in its description but never returned them: the getDataProduct
query did not request members and the Assets field was left empty.

Fetch members via the DataProduct.entities GraphQL resolver in a separate
best-effort query (GetDataProductEntitiesQuery). Issuing it separately
means a DataHub instance whose schema lacks the resolver returns the
product without members rather than failing the whole lookup. The
entities argument is SearchAcrossEntitiesInput, whose query field is
required, so "*" matches all members; the new query is covered by the
schema-validation test against the vendored DataHub schema.

Upgrade DataProduct.Assets from []string to []Entity (urn, name, type)
to match the already-documented output contract (tools-api.md declared
[]Entity; the tools.md example showed objects) and update the
get_data_product output JSON Schema accordingly.

Also pin toolchain go1.26.4 to pick up the patched standard library
(clears reachable govulncheck GO-2026-5039 and GO-2026-5037), and handle
the strings.Builder write returns in rest.go that the linter flagged.

v1.8.1

Toggle v1.8.1's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
fix: return empty entities array instead of null on zero search resul…

…ts (#131)

* fix: return empty entities array instead of null on zero search results

Initialize SearchResult.Entities with make([]SearchEntity, 0, ...) so
json.Marshal produces "entities": [] instead of "entities": null when
there are no results. The OutputSchema declares entities as "type":
"array" (no null allowed), so nil slices caused validation errors.

Fixed in both doSearchAcrossEntities (keyword + semantic) and Search
(legacy client method). Added regression test asserting [] not null.

* test: add client-level regression test for nil entities fix

Add TestSearchAcrossEntities_ZeroResults_EntitiesNotNil that hits a real
httptest server returning zero results and asserts Entities is non-nil
and marshals to [] not null. This exercises the actual make() fix in
doSearchAcrossEntities rather than relying on the mock to pre-initialize
the slice.

v1.8.0

Toggle v1.8.0's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
feat: upgrade datahub_search to use searchAcrossEntities with filters (

…#130)

* feat: upgrade datahub_search to use searchAcrossEntities with filters (#129)

Replace the type-scoped `search` GraphQL query with `searchAcrossEntities`
for keyword mode, enabling advanced field-level filtering and multi-type
search while maintaining full backward compatibility.

New SearchInput fields:
- `types`: search across multiple entity types (overrides entity_type)
- `filters`: advanced field-level filters (fieldPaths, fieldTags, platform,
  owners, domains, glossaryTerms, etc.) that are AND'd together

New client API:
- `SearchFilter` type with Field, Values, Condition, Negated
- `WithTypes()` and `WithOrFilters()` search options
- `SearchAcrossEntities()` client method with full entity fragments

* fix: address review findings for searchAcrossEntities PR

- SearchAcrossEntities now falls back to entityType when types is empty
- SemanticSearch now supports WithTypes and WithSearchFilters options
- SemanticSearchQuery gains DataFlow, Tag, Document entity fragments
  (parity with SearchAcrossEntitiesQuery and SearchQuery)
- Rename WithOrFilters to WithSearchFilters (clearer: filters are AND'd)
- convertFilters merges both Value and Values instead of dropping Value
- Add validateFilters rejecting empty Field or empty Values
- Add tests for all fixes: entityType fallback, types override,
  semantic+filters, Value+Values merge, validation errors

* refactor: extract shared search helper, eliminate query duplication

- Extract doSearchAcrossEntities shared helper; SearchAcrossEntities and
  SemanticSearch are now one-liner delegates (only difference: fulltext flag)
- SemanticSearchQuery is now an alias for SearchAcrossEntitiesQuery —
  single GraphQL query constant with all entity fragments
- Add DefaultEntityType constant to satisfy goconst lint
- Deduplicate values in convertFilters when Value is already in Values
- Add test for client no-types path (all-type search sends no types key)
- Add test for Value+Values deduplication
- Update CLAUDE.md datahub_search description for new capabilities