SAGE LAKEHOUSE BY J79 • FINANCIAL INTELLIGENCE AT COMPUTE SCALE
JJ79Talk to data engineering

Apache Iceberg lakehouse

Ask the data.
Keep it open.

Query normalized Sage financial intelligence through Apache Iceberg storage and the Iceberg REST Catalog from AWS, Azure, or Google Cloud—using the engine your team already operates.

PYSPARK · ICEBERG RESTJ79 CATALOG
# One catalog, open across clouds and engines
spark.sql.catalog.j79 = "org.apache.iceberg.spark.SparkCatalog"
spark.sql.catalog.j79.type = "rest"
spark.sql.catalog.j79.uri = "https://catalog.j79.example/v1"

filings = spark.table("j79.sec.filing_changes")

signals = (
  filings
    .where("form in ('10-K', '10-Q')")
    .where("materiality_score >= 0.80")
    .join(
      spark.table("j79.market.securities"),
      "issuer_id"
    )
)

display(signals.orderBy("filed_at"))
Apache IcebergOpen tables, snapshots, and schema evolution
REST CatalogOne governed discovery and access plane
AWS · Azure · GCPUse object storage and compute in your cloud
Multi-engineSpark, Databricks, Trino, Flink, and more

Ask the Sage lakehouse

From investment objective to evidence-backed allocation.

Natural-language portfolio questions become governed queries across point-in-time fundamentals, J79 Security Scores, risk factors, market histories, and portfolio analytics.

Illustrative model output, not personalized investment advice. Return targets are assumptions, not guarantees.

J79 QUERY PLANNERREADY
01 · Parse objective02 · Query Iceberg snapshots03 · Optimize allocation04 · Model scenarios

Run the question to build an illustrative allocation.

“Diversification is protection against ignorance.”— Warren Buffett

The productive tension

Concentration can create wealth. Diversification can help preserve it.

Famous investors often built fortunes through concentrated conviction. J79’s job is not to maximize the number of holdings—it is to make each concentration explicit, evidence-backed, and appropriate for the investor’s risk budget.

One layer, many workloads

Move from notebook to production without changing the data contract.

Explore filings interactively, engineer factors in batch, join portfolio positions, and operationalize models against stable, documented tables.

01 / RESEARCH

Interactive exploration

Use notebooks and SQL warehouses to examine filing deltas, ownership changes, financial facts, and price context.

02 / MODELS

Feature engineering

Build reproducible training sets from point-in-time facts, disclosure signals, institutional positioning, and market histories.

03 / PORTFOLIOS

Exposure analytics

Join J79 identifiers to internal positions and calculate issuer, sector, factor, filing-event, and liquidity concentrations.

04 / PIPELINES

Scheduled production jobs

Run incremental transforms against new partitions while preserving schema and observation timestamps.

05 / QUALITY

Data expectations

Monitor freshness, row counts, keys, null behavior, and schema revisions before downstream publication.

06 / LINEAGE

Evidence back to source

Carry accession numbers, reporting periods, data provenance, and derivation metadata into downstream outputs.

Fits the lakehouse you operate.

Connection modelsCloud object-store delivery, shared tables, or controlled API hydration depending on security requirements.
Data organizationIssuer reference, filings, structured fundamentals, ownership, transactions, market data, and derived signals.
Incremental processingAs-of partitions and stable keys support repeatable merges and change-data workflows.
Point-in-time analysisObservation and publication timestamps reduce look-ahead bias in research and model evaluation.
PortabilityOpen formats keep workloads usable across Databricks, managed Spark, and self-operated clusters.

Make primary-source intelligence a native lakehouse asset.

Discuss connectivity