{"id":272,"date":"2026-08-23T15:40:14","date_gmt":"2026-08-23T15:40:14","guid":{"rendered":"https:\/\/tobolist.com\/?page_id=272"},"modified":"2026-08-30T19:41:52","modified_gmt":"2026-08-30T19:41:52","slug":"insights","status":"publish","type":"page","link":"https:\/\/tobolist.com\/es\/insights\/","title":{"rendered":"Perceptions"},"content":{"rendered":"\n<div class=\"tb-next tb-page\">\n\n  <!-- ============================================================\n       BARRA DE MARCA\n       ============================================================ -->\n  <nav class=\"tb-nav\" aria-label=\"Main\">\n    <a class=\"tb-nav-brand\" href=\"\/\" aria-label=\"Tobolist - home\">\n      <svg viewBox=\"0 0 402 58\" role=\"img\" aria-hidden=\"true\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\">\n      <path fill=\"currentColor\" d=\"M 203.97 1.55 C 204.73 3.77 205.5 5.96 206.26 8.13 C 207.02 10.3 207.76 12.55 208.47 14.88 C 208.06 14.98 207.68 15.11 207.32 15.26 C 207.02 15.42 206.66 15.55 206.26 15.65 C 205.29 13.48 204.35 11.72 203.43 10.38 C 202.57 8.99 201.7 7.88 200.84 7.05 C 199.72 6.02 198.4 5.32 196.87 4.96 C 195.34 4.6 193.84 4.41 192.37 4.41 C 189.42 4.41 187.48 4.65 186.57 5.11 C 185.55 5.58 185.04 6.51 185.04 7.9 L 185.04 49.04 C 185.04 50.54 185.6 51.59 186.72 52.21 C 187.74 52.78 189.95 53.07 193.36 53.07 L 193.36 55.55 L 171.84 55.55 L 171.84 53.07 C 174.89 53.07 176.95 52.81 178.02 52.29 C 179.45 51.62 180.16 50.54 180.16 49.04 L 180.16 7.28 C 180.16 6.2 179.75 5.45 178.94 5.04 C 178.18 4.62 176.62 4.41 174.28 4.41 C 172.5 4.41 170.67 4.57 168.79 4.88 C 166.96 5.19 165.38 5.91 164.05 7.05 C 163.5 7.52 162.96 8.08 162.45 8.75 C 161.95 9.38 161.36 10.2 160.7 11.23 C 160.29 11.91 159.91 12.57 159.55 13.25 C 159.2 13.87 158.82 14.54 158.41 15.26 C 158.11 15.11 157.8 14.98 157.49 14.88 C 157.24 14.72 156.96 14.56 156.65 14.41 C 157.42 12.24 158.18 10.1 158.94 7.98 C 159.71 5.86 160.47 3.72 161.23 1.55 Z M 203.97 1.55 \"><\/path>\n      <path fill=\"currentColor\" d=\"M 234.5 37.96 C 234.5 43.74 232.88 48.37 229.62 51.83 C 226.52 55.29 222.55 57.02 217.71 57.02 C 212.83 57.02 208.84 55.29 205.73 51.83 C 202.48 48.37 200.85 43.74 200.85 37.96 C 200.85 32.18 202.48 27.55 205.73 24.09 C 208.84 20.63 212.83 18.9 217.71 18.9 C 222.55 18.9 226.52 20.63 229.62 24.09 C 232.88 27.55 234.5 32.18 234.5 37.96 Z M 229.31 37.96 C 229.31 33.21 228.19 29.31 225.95 26.26 C 223.87 23.21 221.12 21.69 217.71 21.69 C 214.36 21.69 211.55 23.21 209.32 26.26 C 207.13 29.31 206.04 33.21 206.04 37.96 C 206.04 42.66 207.11 46.56 209.24 49.66 C 211.43 52.65 214.25 54.15 217.71 54.15 C 221.02 54.15 223.77 52.65 225.95 49.66 C 228.19 46.61 229.31 42.71 229.31 37.96 Z M 229.31 37.96 \"><\/path>\n      <path fill=\"currentColor\" d=\"M 269.72 37.8 C 269.72 43.43 268.14 48 264.99 51.52 C 263.52 53.17 261.89 54.46 260.11 55.39 C 258.38 56.37 256.37 56.86 254.08 56.86 C 249.65 56.86 246.04 55.52 243.24 52.83 C 242.83 53.35 242.45 53.84 242.09 54.3 C 241.79 54.82 241.43 55.34 241.03 55.86 L 239.27 55.86 L 239.27 9.3 C 239.27 8.16 239.02 7.36 238.51 6.89 C 238 6.38 236.78 6.12 234.84 6.12 C 234.79 5.86 234.74 5.58 234.69 5.27 C 234.64 4.96 234.59 4.68 234.54 4.41 C 236.17 3.8 237.75 3.18 239.27 2.55 C 240.8 1.94 242.35 1.32 243.93 0.7 L 243.93 22 C 246.42 19.93 249.8 18.9 254.08 18.9 C 258.76 18.9 262.52 20.58 265.37 23.94 C 268.27 27.35 269.72 31.97 269.72 37.8 Z M 264.61 37.88 C 264.61 33.18 263.52 29.39 261.33 26.5 C 260.26 25.1 258.96 24.02 257.43 23.24 C 255.96 22.46 254.28 22.08 252.4 22.08 C 249.14 22.08 246.32 23.27 243.93 25.64 L 243.93 49.81 C 246.27 52.34 249.09 53.61 252.4 53.61 C 255.8 53.61 258.71 52.09 261.1 49.04 C 263.44 46.2 264.61 42.48 264.61 37.88 Z M 264.61 37.88 \"><\/path>\n      <path fill=\"currentColor\" d=\"M 308.05 37.96 C 308.05 43.74 306.42 48.37 303.16 51.83 C 300.06 55.29 296.09 57.02 291.26 57.02 C 286.37 57.02 282.38 55.29 279.27 51.83 C 276.02 48.37 274.39 43.74 274.39 37.96 C 274.39 32.18 276.02 27.55 279.27 24.09 C 282.38 20.63 286.37 18.9 291.26 18.9 C 296.09 18.9 300.06 20.63 303.16 24.09 C 306.42 27.55 308.05 32.18 308.05 37.96 Z M 302.86 37.96 C 302.86 33.21 301.74 29.31 299.5 26.26 C 297.41 23.21 294.66 21.69 291.26 21.69 C 287.9 21.69 285.1 23.21 282.86 26.26 C 280.68 29.31 279.58 33.21 279.58 37.96 C 279.58 42.66 280.65 46.56 282.79 49.66 C 284.97 52.65 287.8 54.15 291.26 54.15 C 294.56 54.15 297.31 52.65 299.5 49.66 C 301.74 46.61 302.86 42.71 302.86 37.96 Z M 302.86 37.96 \"><\/path>\n      <path fill=\"currentColor\" d=\"M 324.46 55.55 L 309.81 55.55 L 309.81 53.07 C 311.8 53.07 313.14 52.78 313.86 52.21 C 314.46 51.7 314.77 50.59 314.77 48.88 L 314.77 8.52 C 314.77 7.64 314.46 7.02 313.86 6.66 C 313.25 6.25 311.95 6.04 309.96 6.04 L 309.51 4.18 C 311.13 3.56 312.79 2.97 314.46 2.4 C 316.14 1.78 317.8 1.16 319.43 0.54 L 319.43 48.88 C 319.43 50.48 319.71 51.57 320.27 52.14 C 320.67 52.39 321.21 52.6 321.87 52.76 C 322.58 52.96 323.45 53.07 324.46 53.07 Z M 324.46 55.55 \"><\/path>\n      <path fill=\"currentColor\" d=\"M 340.23 4.34 C 340.23 5.58 339.82 6.61 339.01 7.44 C 338.19 8.26 337.18 8.68 335.95 8.68 C 334.84 8.68 333.89 8.26 333.13 7.44 C 332.42 6.61 332.06 5.58 332.06 4.34 C 332.06 3.1 332.44 2.07 333.21 1.24 C 333.97 0.41 334.89 0 335.95 0 C 337.18 0 338.19 0.41 339.01 1.24 C 339.82 2.07 340.23 3.1 340.23 4.34 Z M 343.81 55.62 L 329.54 55.62 L 329.54 53.07 C 331.63 53.07 332.98 52.76 333.59 52.14 C 334.1 51.67 334.35 50.82 334.35 49.58 L 334.35 27.27 C 334.35 26.13 334.12 25.36 333.66 24.95 C 333.11 24.48 331.76 24.25 329.62 24.25 L 329.31 22.54 L 339.01 18.82 L 339.01 49.58 C 339.01 51.03 339.23 51.93 339.69 52.29 C 340.36 52.81 341.73 53.07 343.81 53.07 Z M 343.81 55.62 \"><\/path>\n      <path fill=\"currentColor\" d=\"M 375.11 46.64 C 375.11 48.75 374.52 50.64 373.36 52.29 C 373.1 52.55 372.84 52.81 372.59 53.07 C 372.39 53.32 372.16 53.58 371.9 53.84 C 369.61 55.86 366.51 56.86 362.59 56.86 C 359.9 56.86 357.56 56.42 355.57 55.55 C 355.32 55.44 355.06 55.36 354.81 55.31 C 354.61 55.31 354.38 55.31 354.12 55.31 L 353.36 55.31 C 352.8 55.57 352.27 55.78 351.76 55.93 C 351.25 56.14 350.71 56.34 350.16 56.55 C 349.44 54.59 348.7 52.68 347.94 50.82 C 347.18 48.91 346.44 46.97 345.73 45.01 C 345.93 44.86 346.11 44.7 346.26 44.54 C 346.46 44.39 346.7 44.26 346.95 44.16 C 349.14 47.26 351.22 49.53 353.21 50.97 C 356.36 53.04 359.46 54.07 362.52 54.07 C 364.91 54.07 366.92 53.32 368.55 51.83 C 369.87 50.48 370.53 48.86 370.53 46.95 C 370.53 45.71 370.05 44.6 369.08 43.61 C 368.88 43.46 368.62 43.28 368.32 43.07 C 368.01 42.87 367.66 42.63 367.25 42.38 C 366.59 41.96 365.72 41.55 364.66 41.14 C 363.59 40.67 362.29 40.18 360.76 39.66 C 358.37 38.84 356.39 38.09 354.81 37.42 C 353.23 36.7 352.04 36 351.22 35.32 C 348.17 33.41 346.64 30.88 346.64 27.73 C 346.64 24.84 347.99 22.54 350.69 20.84 C 352.72 19.55 355.27 18.9 358.32 18.9 C 360.96 18.9 363.69 19.42 366.48 20.45 C 366.89 20.35 367.32 20.17 367.79 19.91 C 368.24 19.6 368.75 19.24 369.31 18.82 C 370.02 20.68 370.73 22.57 371.45 24.48 C 372.16 26.34 372.87 28.22 373.58 30.14 C 373.33 30.24 373.07 30.37 372.82 30.52 C 372.62 30.68 372.39 30.81 372.13 30.91 C 370.86 28.43 369.16 26.42 367.02 24.87 C 363.87 22.75 360.96 21.69 358.32 21.69 C 356.39 21.69 354.86 22.03 353.74 22.7 C 352.06 23.63 351.22 25.23 351.22 27.5 C 351.22 28.95 351.96 30.19 353.44 31.22 C 353.64 31.32 353.84 31.43 354.05 31.53 C 354.25 31.63 354.48 31.74 354.73 31.84 C 355.45 32.25 356.34 32.69 357.41 33.16 C 358.47 33.62 359.67 34.06 360.99 34.47 C 363.74 35.35 366 36.23 367.79 37.11 C 369.61 37.93 371.02 38.71 371.98 39.43 C 374.07 41.14 375.11 43.54 375.11 46.64 Z M 375.11 46.64 \"><\/path>\n      <path fill=\"currentColor\" d=\"M 396.65 47.95 C 396.91 48.16 397.14 48.34 397.34 48.5 C 397.59 48.65 397.85 48.8 398.10 48.96 C 397.19 51.13 395.94 52.94 394.36 54.38 C 392.58 56.09 390.65 56.94 388.56 56.94 C 385.82 56.94 383.73 55.93 382.3 53.92 C 381.9 53.2 381.57 52.37 381.31 51.44 C 381.06 50.51 380.93 49.48 380.93 48.34 L 380.93 22.7 L 374.98 22.7 L 374.98 20.22 C 377.73 20.22 379.76 19.16 381.08 17.04 C 382.56 14.62 383.3 11.16 383.3 6.66 L 385.59 6.66 L 385.59 19.68 L 395.58 19.68 L 395.58 22.7 L 385.59 22.7 L 385.59 48.88 C 385.59 51.62 387.01 52.99 389.86 52.99 C 392.15 52.99 394.41 51.31 396.65 47.95 Z M 396.65 47.95 \"><\/path>\n      <path fill=\"var(--tb-mark,#FF4001)\" d=\"M 102.74 0.61 L 30.22 0.61 C 15.14 0.61 2.91 13 2.91 28.29 L 2.91 29.11 C 2.91 44.4 15.14 56.8 30.22 56.8 L 102.74 56.8 C 117.82 56.8 130.05 44.4 130.05 29.11 L 130.05 28.29 C 130.05 13 117.82 0.61 102.74 0.61 Z M 117.73 45.16 C 113.47 49.48 107.8 51.86 101.78 51.86 C 95.75 51.86 90.08 49.48 85.82 45.16 C 84.41 43.72 83.19 42.11 82.21 40.38 C 78.98 34.67 72.98 31.15 66.5 31.14 L 66.48 31.14 C 60 31.14 54.01 34.65 50.77 40.35 C 49.79 42.09 48.56 43.7 47.12 45.16 C 42.86 49.48 37.2 51.86 31.17 51.86 C 25.14 51.86 19.47 49.48 15.21 45.16 C 6.41 36.24 6.41 21.72 15.21 12.8 C 19.48 8.48 25.14 6.11 31.17 6.11 C 37.2 6.11 42.87 8.48 47.12 12.8 C 48.56 14.25 49.79 15.87 50.77 17.61 C 54.01 23.31 60 26.82 66.48 26.82 L 66.5 26.82 C 72.99 26.81 78.98 23.29 82.21 17.59 C 83.19 15.85 84.41 14.24 85.82 12.8 C 90.09 8.48 95.75 6.11 101.78 6.11 C 107.8 6.11 113.48 8.48 117.73 12.8 C 122 17.12 124.34 22.87 124.34 28.98 C 124.34 35.09 122 40.84 117.73 45.16 Z M 117.73 45.16 \"><\/path>\n      <\/svg>\n    <\/a>\n    <ul class=\"tb-nav-links\">\n      <li><a href=\"\/services\/\">Services<\/a><\/li>\n      <li><a href=\"\/cases\/\">Work<\/a><\/li>\n      <li><a href=\"\/insights\/\" aria-current=\"page\">Insights<\/a><\/li>\n      <li><a href=\"\/about\/\">About<\/a><\/li>\n    <\/ul>\n    <a class=\"tb-nav-cta\" href=\"\/contact\/\">Contact<\/a>\n  <\/nav>\n\n  <!-- ============================================================\n       01 \u00b7 APERTURA\n       El posicionamiento est\u00e1 en la entradilla: este glosario est\u00e1\n       escrito para quien tiene que APROBAR un proyecto de datos, no\n       para quien lo construye. Casi todos los glosarios del sector\n       est\u00e1n escritos al rev\u00e9s, y por eso no los lee nadie que decida.\n       ============================================================ -->\n  <header class=\"tb-sec\">\n    <p class=\"tb-eyebrow-2\">Insights<\/p>\n    <h1 class=\"tb-ins-h1\">The data glossary, <em>written for the person paying for it<\/em>.<\/h1>\n    <p class=\"tb-lede tb-ins-open-lede\">\n      Every data proposal you receive will be full of names: Databricks, Snowflake, Airflow,\n      Fabric, lakehouse, streaming. Most glossaries explain them to engineers. This one\n      explains them to whoever has to decide, approve or defend the project \u2014 including\n      which ones are genuinely different and which ones are the same idea sold by a\n      different vendor.\n    <\/p>\n    <p class=\"tb-ins-open-meta\">\n      <span><b>29<\/b> terms<\/span>\n      <span><b>6<\/b> sections<\/span>\n      <span>Free to use and to quote<\/span>\n      <span>No sign-up, no PDF<\/span>\n    <\/p>\n  <\/header>\n\n\n  <!-- ============================================================\n       02 \u00b7 \u00cdNDICE DE T\u00c9RMINOS\n       Chips cuadrados enlazados a las anclas. Adem\u00e1s de ser \u00fatil,\n       es la parte que hace que un motor entienda de un vistazo que\n       esta URL cubre 29 conceptos y no uno.\n       ============================================================ -->\n  <section class=\"tb-sec\" id=\"index\">\n    <div class=\"tb-ins-col\">\n      <p class=\"tb-lead-label\">Index<\/p>\n      <h2 class=\"tb-h2\">Jump straight to <em>the word you were sent.<\/em><\/h2>\n        <p class=\"tb-ins-idx-k\">01 \u00b7 The three roles<\/p>\n        <ul class=\"tb-ins-index\">\n          <li><a class=\"tb-ins-chip\" href=\"#data-architect\">Data architect<\/a><\/li>\n          <li><a class=\"tb-ins-chip\" href=\"#data-engineer\">Data engineer<\/a><\/li>\n          <li><a class=\"tb-ins-chip\" href=\"#bi-analyst\">BI analyst<\/a><\/li>\n        <\/ul>\n        <p class=\"tb-ins-idx-k\">02 \u00b7 Platforms and processing<\/p>\n        <ul class=\"tb-ins-index\">\n          <li><a class=\"tb-ins-chip\" href=\"#databricks\">Databricks<\/a><\/li>\n          <li><a class=\"tb-ins-chip\" href=\"#apache-spark\">Apache Spark<\/a><\/li>\n          <li><a class=\"tb-ins-chip\" href=\"#snowflake\">Snowflake<\/a><\/li>\n          <li><a class=\"tb-ins-chip\" href=\"#google-bigquery\">Google BigQuery<\/a><\/li>\n          <li><a class=\"tb-ins-chip\" href=\"#azure-synapse\">Azure Synapse Analytics<\/a><\/li>\n          <li><a class=\"tb-ins-chip\" href=\"#microsoft-fabric\">Microsoft Fabric<\/a><\/li>\n          <li><a class=\"tb-ins-chip\" href=\"#amazon-redshift\">Amazon Redshift<\/a><\/li>\n        <\/ul>\n        <p class=\"tb-ins-idx-k\">03 \u00b7 Where data is stored<\/p>\n        <ul class=\"tb-ins-index\">\n          <li><a class=\"tb-ins-chip\" href=\"#data-lake\">Data lake<\/a><\/li>\n          <li><a class=\"tb-ins-chip\" href=\"#data-warehouse\">Data warehouse<\/a><\/li>\n          <li><a class=\"tb-ins-chip\" href=\"#lakehouse\">Lakehouse<\/a><\/li>\n          <li><a class=\"tb-ins-chip\" href=\"#amazon-s3\">Amazon S3<\/a><\/li>\n          <li><a class=\"tb-ins-chip\" href=\"#google-cloud-storage\">Google Cloud Storage<\/a><\/li>\n          <li><a class=\"tb-ins-chip\" href=\"#azure-blob-storage\">Azure Blob Storage \/ ADLS<\/a><\/li>\n        <\/ul>\n        <p class=\"tb-ins-idx-k\">04 \u00b7 Moving and orchestrating<\/p>\n        <ul class=\"tb-ins-index\">\n          <li><a class=\"tb-ins-chip\" href=\"#data-pipeline\">Data pipeline<\/a><\/li>\n          <li><a class=\"tb-ins-chip\" href=\"#etl-elt\">ETL \/ ELT<\/a><\/li>\n          <li><a class=\"tb-ins-chip\" href=\"#apache-airflow\">Apache Airflow<\/a><\/li>\n          <li><a class=\"tb-ins-chip\" href=\"#cloud-composer\">Google Cloud Composer<\/a><\/li>\n          <li><a class=\"tb-ins-chip\" href=\"#azure-data-factory\">Azure Data Factory<\/a><\/li>\n          <li><a class=\"tb-ins-chip\" href=\"#dbt\">dbt<\/a><\/li>\n        <\/ul>\n        <p class=\"tb-ins-idx-k\">05 \u00b7 Real time<\/p>\n        <ul class=\"tb-ins-index\">\n          <li><a class=\"tb-ins-chip\" href=\"#batch-vs-streaming\">Batch vs streaming<\/a><\/li>\n          <li><a class=\"tb-ins-chip\" href=\"#apache-kafka\">Apache Kafka<\/a><\/li>\n          <li><a class=\"tb-ins-chip\" href=\"#pubsub\">Google Cloud Pub\/Sub<\/a><\/li>\n        <\/ul>\n        <p class=\"tb-ins-idx-k\">06 \u00b7 Reporting and trust<\/p>\n        <ul class=\"tb-ins-index\">\n          <li><a class=\"tb-ins-chip\" href=\"#power-bi\">Power BI<\/a><\/li>\n          <li><a class=\"tb-ins-chip\" href=\"#kpi\">KPI<\/a><\/li>\n          <li><a class=\"tb-ins-chip\" href=\"#semantic-model\">Semantic model<\/a><\/li>\n          <li><a class=\"tb-ins-chip\" href=\"#data-quality\">Data quality<\/a><\/li>\n        <\/ul>\n    <\/div>\n  <\/section>\n\n  <!-- ============================================================\n       GRUPO 01 \u00b7 The three roles\n       ============================================================ -->\n  <section class=\"tb-sec tb-ins-group\" id=\"roles\" data-n=\"01\">\n    <div class=\"tb-ins-col\">\n      <p class=\"tb-lead-label\">01 \u00b7 The three roles<\/p>\n      <h2 class=\"tb-h2\">Who does what, <em>and in what order.<\/em><\/h2>\n      <p class=\"tb-lede\">Almost every data project touches all three. Confusing them is the most common reason a project is scoped wrong: you hire someone to build a pipeline when the real problem was that nobody decided what the platform was for.<\/p>\n\n      <article class=\"tb-ins-term\" id=\"data-architect\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Data architect<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Role<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">Decides <strong>how the data platform should be built<\/strong>, before anyone builds it. Chooses the technologies and architecture patterns, defines how data is stored, processed and integrated, and sets the standards on scalability, security and governance that the engineering team then follows.<\/p>\n        <p class=\"tb-ins-note\"><b>Not the same as a data engineer<\/b>An architect works at platform level and mostly produces decisions and designs. An engineer works at system level and produces things that run. On a small project one person does both; on a corporate platform, treating them as the same job is how you end up with three incompatible pipelines.<\/p>\n        <p class=\"tb-ins-see\"><a href=\"\/services\/data-architecture\/\">Data architecture service <i>\u2192<\/i><\/a><\/p>\n      <\/article>\n\n      <article class=\"tb-ins-term\" id=\"data-engineer\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Data engineer<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Role<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">Builds and maintains the infrastructure that moves, transforms, stores and processes data. The job is to get data out of the systems where it lives \u2014 ERP, CRM, internal apps, databases \u2014 and into the place where it will actually be used, reliably and on schedule.<\/p>\n        <p class=\"tb-ins-d\">Typical stack: Python, SQL, Spark, Databricks, Kafka, Airflow, dbt, and one of the three big clouds.<\/p>\n        <p class=\"tb-ins-see\"><a href=\"\/services\/data-engineering\/\">Data engineering service <i>\u2192<\/i><\/a><\/p>\n      <\/article>\n\n      <article class=\"tb-ins-term\" id=\"bi-analyst\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">BI analyst<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Role<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">Turns available data into something a business can decide with: KPIs, dashboards, reports, reporting-oriented data models. Sits closest to the business user of the three roles \u2014 the job is not to build the platform but to make the numbers legible and defensible.<\/p>\n        <p class=\"tb-ins-d\">Typical stack: Power BI, Tableau, Looker, SQL, Excel, Microsoft Fabric.<\/p>\n        <p class=\"tb-ins-see\"><a href=\"\/services\/business-intelligence\/\">Business intelligence service <i>\u2192<\/i><\/a><\/p>\n      <\/article>\n\n      <p class=\"tb-ins-up\"><a href=\"#index\">\u2191 Back to the index<\/a><\/p>\n    <\/div>\n  <\/section>\n\n\n  <!-- ============================================================\n       GRUPO 02 \u00b7 Platforms and processing\n       ============================================================ -->\n  <section class=\"tb-sec tb-ins-group\" id=\"platforms\" data-n=\"02\">\n    <div class=\"tb-ins-col\">\n      <p class=\"tb-lead-label\">02 \u00b7 Platforms and processing<\/p>\n      <h2 class=\"tb-h2\">Where the heavy work <em>actually happens.<\/em><\/h2>\n      <p class=\"tb-lede\">These are the names that come up first in any conversation about a modern data platform. They overlap, and that overlap is exactly what makes them easy to confuse.<\/p>\n\n      <article class=\"tb-ins-term\" id=\"databricks\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Databricks<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Processing platform<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">A platform for processing, engineering and analysing large volumes of data, built around Apache Spark. Different profiles \u2014 engineers, analysts, data scientists \u2014 work on the same platform to transform data, run jobs, train machine learning models and build data products.<\/p>\n        <p class=\"tb-ins-d\">It is <strong>not tied to one cloud<\/strong>: it runs on Azure, AWS and Google Cloud alike.<\/p>\n        <p class=\"tb-ins-note\"><b>Not the same as Snowflake<\/b>Databricks leans towards processing and engineering; Snowflake leans towards storing and querying structured data for analytics. They compete in the middle and plenty of companies run both. Anyone who tells you they are interchangeable is selling one of them.<\/p>\n        <p class=\"tb-ins-see\"><a href=\"\/case-industrial\/\">A Databricks warehouse on SAP data <i>\u2192<\/i><\/a><\/p>\n      <\/article>\n\n      <article class=\"tb-ins-term\" id=\"apache-spark\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Apache Spark<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Processing engine<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">The open-source engine for distributed data processing that sits underneath a large part of the modern data stack, Databricks included. When a job is too big for one machine, Spark splits it across many. PySpark is its Python interface \u2014 the one most data engineers actually write.<\/p>\n        <p class=\"tb-ins-see\"><a href=\"\/case-telco\/\">PySpark in a real-time telco pipeline <i>\u2192<\/i><\/a><\/p>\n      <\/article>\n\n      <article class=\"tb-ins-term\" id=\"snowflake\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Snowflake<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Cloud data warehouse<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">A cloud platform specialised in storing, processing and querying large amounts of data for analytics, reporting, BI and data science. Like Databricks, it is <strong>independent of any single cloud<\/strong>: it runs on top of AWS, Azure or Google Cloud infrastructure rather than belonging to any of them.<\/p>\n      <\/article>\n\n      <article class=\"tb-ins-term\" id=\"google-bigquery\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Google BigQuery<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Cloud data warehouse<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">Google Cloud&#8217;s managed data warehouse and analytics platform. Stores very large volumes of data and runs queries over them without you managing servers. If a company has chosen Google Cloud, BigQuery is usually the centre of its analytics.<\/p>\n      <\/article>\n\n      <article class=\"tb-ins-term\" id=\"azure-synapse\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Azure Synapse Analytics<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Analytics platform<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">Microsoft Azure&#8217;s analytics and data processing platform, combining storage, processing and analysis. For years it was the centre of Microsoft&#8217;s data ecosystem. Microsoft is now pushing <strong>Microsoft Fabric<\/strong>, so new Azure projects increasingly mention Fabric instead.<\/p>\n        <p class=\"tb-ins-note\"><b>Hearing \u201cSynapse\u201d tells you the cloud, not the product<\/b>In practice, when Synapse comes up in a conversation the most useful thing it tells you is that the project lives on Microsoft Azure. What the client actually needs may end up being Fabric, Databricks or something else entirely.<\/p>\n      <\/article>\n\n      <article class=\"tb-ins-term\" id=\"microsoft-fabric\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Microsoft Fabric<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Analytics platform<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">Microsoft&#8217;s unified analytics platform, bringing storage, data engineering, warehousing and Power BI reporting into one product. It is where Microsoft is steering new Azure data projects, and it is increasingly the default answer inside Microsoft-heavy organisations.<\/p>\n      <\/article>\n\n      <article class=\"tb-ins-term\" id=\"amazon-redshift\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Amazon Redshift<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Cloud data warehouse<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">AWS&#8217;s managed data warehouse. Plays the role that BigQuery plays on Google Cloud and that Synapse or Fabric play on Azure: the place where structured data is stored and queried for analytics.<\/p>\n      <\/article>\n\n      <p class=\"tb-ins-up\"><a href=\"#index\">\u2191 Back to the index<\/a><\/p>\n    <\/div>\n  <\/section>\n\n\n  <!-- ============================================================\n       GRUPO 03 \u00b7 Where data is stored\n       ============================================================ -->\n  <section class=\"tb-sec tb-ins-group\" id=\"storage\" data-n=\"03\">\n    <div class=\"tb-ins-col\">\n      <p class=\"tb-lead-label\">03 \u00b7 Where data is stored<\/p>\n      <h2 class=\"tb-h2\">Lake, warehouse, <em>lakehouse.<\/em><\/h2>\n      <p class=\"tb-lede\">Three words that get used interchangeably in sales conversations and mean genuinely different things. Getting them straight is the fastest way to tell whether a proposal has been thought through.<\/p>\n\n      <article class=\"tb-ins-term\" id=\"data-lake\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Data lake<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Storage pattern<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">A repository that holds data <strong>as it arrives<\/strong>, in whatever shape it arrives: files, logs, exports, images, raw dumps. Cheap and flexible, because nothing has to be modelled before it lands. The cost is that a lake with no governance quietly becomes a place where data goes to be forgotten.<\/p>\n      <\/article>\n\n      <article class=\"tb-ins-term\" id=\"data-warehouse\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Data warehouse<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Storage pattern<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">A repository that holds <strong>structured, modelled data<\/strong> ready for analysis: sales by customer, monthly billing, margins by region. Data has to be cleaned and shaped before it goes in, which is more work up front and far less work every time somebody asks a question.<\/p>\n      <\/article>\n\n      <article class=\"tb-ins-term\" id=\"lakehouse\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Lakehouse<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Storage pattern<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">An architecture that tries to keep the cheap, flexible storage of a lake and add the structure, reliability and query performance of a warehouse on top of it. It is the pattern behind most modern platform designs, and the one Databricks built its product around.<\/p>\n      <\/article>\n\n      <article class=\"tb-ins-term\" id=\"amazon-s3\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Amazon S3<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Object storage \u00b7 AWS<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">AWS&#8217;s object storage service, and one of the foundations of most data architectures built on AWS: files, datasets, logs, backups, raw data, processing output. Very often it <em>is<\/em> the data lake.<\/p>\n      <\/article>\n\n      <article class=\"tb-ins-term\" id=\"google-cloud-storage\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Google Cloud Storage<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Object storage \u00b7 GCP<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">Google Cloud&#8217;s object storage service, usually shortened to GCS. Does essentially the same job as Amazon S3 within the Google Cloud ecosystem.<\/p>\n      <\/article>\n\n      <article class=\"tb-ins-term\" id=\"azure-blob-storage\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Azure Blob Storage \/ ADLS<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Object storage \u00b7 Azure<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">Microsoft Azure&#8217;s object storage. ADLS \u2014 Azure Data Lake Storage \u2014 is the variant built for analytics workloads. Together they are the Azure equivalent of S3 or GCS.<\/p>\n      <\/article>\n\n      <p class=\"tb-ins-up\"><a href=\"#index\">\u2191 Back to the index<\/a><\/p>\n    <\/div>\n  <\/section>\n\n\n  <!-- ============================================================\n       GRUPO 04 \u00b7 Moving and orchestrating\n       ============================================================ -->\n  <section class=\"tb-sec tb-ins-group\" id=\"moving\" data-n=\"04\">\n    <div class=\"tb-ins-col\">\n      <p class=\"tb-lead-label\">04 \u00b7 Moving and orchestrating<\/p>\n      <h2 class=\"tb-h2\">Getting data from <em>there to here,<\/em> on time.<\/h2>\n      <p class=\"tb-lede\">Most of the effort in a data project is not analysis. It is moving data reliably, in the right order, and knowing within minutes when something has failed.<\/p>\n\n      <article class=\"tb-ins-term\" id=\"data-pipeline\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Data pipeline<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Concept<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">The set of automated steps that take data from where it is produced to where it is consumed: extract it, clean it, transform it, load it, and do it again tomorrow without anybody pressing a button. When people say a data project failed, they usually mean the pipeline stopped being trustworthy.<\/p>\n      <\/article>\n\n      <article class=\"tb-ins-term\" id=\"etl-elt\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">ETL \/ ELT<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Concept<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">Two orderings of the same three steps. <strong>ETL<\/strong> \u2014 extract, transform, load \u2014 cleans the data before storing it. <strong>ELT<\/strong> \u2014 extract, load, transform \u2014 stores it raw first and transforms it inside the destination platform, which is what modern cloud warehouses are fast enough to allow.<\/p>\n      <\/article>\n\n      <article class=\"tb-ins-term\" id=\"apache-airflow\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Apache Airflow<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Orchestration<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">The tool that organises, schedules and supervises data processes. Think of a conductor: at 02:00 pull this data, when it finishes run this transformation, then refresh this table, and raise an alert if any step fails. That job is called <strong>pipeline orchestration<\/strong>.<\/p>\n        <p class=\"tb-ins-see\"><a href=\"\/case-telco\/\">Airflow orchestrating PySpark at the edge <i>\u2192<\/i><\/a><\/p>\n      <\/article>\n\n      <article class=\"tb-ins-term\" id=\"cloud-composer\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Google Cloud Composer<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Orchestration \u00b7 GCP<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">Google Cloud&#8217;s managed Airflow service. Composer <em>is<\/em> Airflow, run and maintained by Google.<\/p>\n        <p class=\"tb-ins-note\"><b>Not a different technology<\/b>If a client says \u201cwe&#8217;re on GCP and we use Composer\u201d, you can read it as \u201cthey use Airflow inside Google Cloud\u201d. An engineer who knows Airflow well transfers to Composer with very little friction \u2014 which matters when you&#8217;re deciding whether a stack mismatch is real or cosmetic.<\/p>\n      <\/article>\n\n      <article class=\"tb-ins-term\" id=\"azure-data-factory\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Azure Data Factory<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Integration \u00b7 Azure<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">Microsoft Azure&#8217;s service for integrating, moving and orchestrating data between systems \u2014 ERP, SQL Server, files, APIs, external applications \u2014 and delivering it into the data platform. Usually shortened to ADF, and near-universal in Azure-based companies.<\/p>\n        <p class=\"tb-ins-note\"><b>Overlaps with Airflow, but isn&#8217;t the same<\/b>Both can coordinate pipelines, so they come up in the same conversations. ADF is stronger on connecting to and moving data between systems; Airflow is stronger on orchestrating arbitrary logic. Plenty of Azure platforms run both, each for what it&#8217;s good at.<\/p>\n        <p class=\"tb-ins-see\"><a href=\"\/case-data-quality\/\">ADF jobs feeding a quality platform <i>\u2192<\/i><\/a><\/p>\n      <\/article>\n\n      <article class=\"tb-ins-term\" id=\"dbt\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">dbt<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Transformation<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">A tool for transforming and organising data using mostly SQL. The data has usually already landed in a warehouse, and dbt turns it from raw tables into structured, documented, tested business concepts: sales by customer, monthly billing, margin by region.<\/p>\n        <p class=\"tb-ins-note\"><b>Doesn&#8217;t replace your warehouse<\/b>dbt is not an alternative to Databricks, Snowflake or BigQuery \u2014 it runs on top of them. Snowflake + dbt and BigQuery + dbt are both completely normal combinations.<\/p>\n      <\/article>\n\n      <p class=\"tb-ins-up\"><a href=\"#index\">\u2191 Back to the index<\/a><\/p>\n    <\/div>\n  <\/section>\n\n\n  <!-- ============================================================\n       GRUPO 05 \u00b7 Real time\n       ============================================================ -->\n  <section class=\"tb-sec tb-ins-group\" id=\"realtime\" data-n=\"05\">\n    <div class=\"tb-ins-col\">\n      <p class=\"tb-lead-label\">05 \u00b7 Real time<\/p>\n      <h2 class=\"tb-h2\">Batch or streaming \u2014 <em>and why it matters.<\/em><\/h2>\n      <p class=\"tb-lede\">The single question that changes the cost and the architecture of a project more than any other. It is worth asking early and answering honestly.<\/p>\n\n      <article class=\"tb-ins-term\" id=\"batch-vs-streaming\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Batch vs streaming<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Concept<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\"><strong>Batch<\/strong> means processing data in scheduled blocks: every night we process the day&#8217;s sales. <strong>Streaming<\/strong> means processing events as they happen: we want to see each sale as it occurs. Streaming is more expensive to build and to run, and it is genuinely necessary far less often than it gets asked for.<\/p>\n      <\/article>\n\n      <article class=\"tb-ins-term\" id=\"apache-kafka\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Apache Kafka<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Streaming<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">The best-known technology for carrying large volumes of information continuously and in near real time: transactions, application events, sensor readings, user activity, logs. If Kafka is in the requirements, the project is usually a fairly technical data engineering job.<\/p>\n        <p class=\"tb-ins-see\"><a href=\"\/case-telco\/\">Kafka moving edge data in real time <i>\u2192<\/i><\/a><\/p>\n      <\/article>\n\n      <article class=\"tb-ins-term\" id=\"pubsub\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Google Cloud Pub\/Sub<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Streaming \u00b7 GCP<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">Google Cloud&#8217;s managed service for sending and receiving events or messages between systems, used in real-time and event-driven architectures.<\/p>\n        <p class=\"tb-ins-note\"><b>Conceptually close to Kafka, not identical<\/b>Kafka is a streaming platform you can deploy in many ways, anywhere. Pub\/Sub is a managed Google Cloud service. The concepts transfer well between them; the operational realities do not.<\/p>\n      <\/article>\n\n      <p class=\"tb-ins-up\"><a href=\"#index\">\u2191 Back to the index<\/a><\/p>\n    <\/div>\n  <\/section>\n\n\n  <!-- ============================================================\n       GRUPO 06 \u00b7 Reporting and trust\n       ============================================================ -->\n  <section class=\"tb-sec tb-ins-group\" id=\"reporting\" data-n=\"06\">\n    <div class=\"tb-ins-col\">\n      <p class=\"tb-lead-label\">06 \u00b7 Reporting and trust<\/p>\n      <h2 class=\"tb-h2\">The part the <em>business actually sees.<\/em><\/h2>\n      <p class=\"tb-lede\">Everything upstream exists so that somebody can look at a number and act on it. If they don&#8217;t trust the number, none of the rest counted.<\/p>\n\n      <article class=\"tb-ins-term\" id=\"power-bi\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Power BI<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Reporting<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">Microsoft&#8217;s business intelligence and data visualisation tool, and by some distance the most common reporting layer in European mid-market companies. Connects to data sources, models the data and produces the dashboards a management team reads.<\/p>\n        <p class=\"tb-ins-see\"><a href=\"\/case-finance\/\">SAP accounting data surfaced in Power BI <i>\u2192<\/i><\/a><\/p>\n      <\/article>\n\n      <article class=\"tb-ins-term\" id=\"kpi\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">KPI<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Concept<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">A key performance indicator: a single number that is supposed to tell you whether something is going well. The hard part is almost never calculating it \u2014 it is agreeing on its definition across departments so that finance and sales don&#8217;t produce two different revenue figures from the same data.<\/p>\n      <\/article>\n\n      <article class=\"tb-ins-term\" id=\"semantic-model\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Semantic model<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Concept<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">The layer that defines what each business concept means in the reporting tool: what counts as revenue, what counts as an active customer, how a margin is calculated. It is where a definition lives once instead of being re-invented in every dashboard.<\/p>\n      <\/article>\n\n      <article class=\"tb-ins-term\" id=\"data-quality\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Data quality<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Concept<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">Whether the data is complete, correct, consistent and current enough to be used. Quality problems that are only detected in the reporting layer have already cost you the decision that was made on them, which is why checks belong upstream in the pipeline.<\/p>\n        <p class=\"tb-ins-see\"><a href=\"\/case-data-quality\/\">Quality rules and AI-written reports <i>\u2192<\/i><\/a><\/p>\n      <\/article>\n\n      <p class=\"tb-ins-up\"><a href=\"#index\">\u2191 Back to the index<\/a><\/p>\n    <\/div>\n  <\/section>\n\n  <!-- ============================================================\n       08 \u00b7 TABLA DE EQUIVALENCIAS ENTRE NUBES\n       La pieza m\u00e1s citable de toda la p\u00e1gina: es la respuesta\n       directa a \"\u00bfcu\u00e1l es el equivalente de X en Azure\/AWS\/GCP?\",\n       que es una de las preguntas m\u00e1s frecuentes del sector.\n       << Nombra servicios de AWS y Azure m\u00e1s all\u00e1 de los que\n          aparecen en vuestros casos publicados (MWAA, Kinesis, MSK,\n          Event Hubs). Son datos correctos del sector y la tabla es\n          una referencia, no una declaraci\u00f3n de experiencia. Si aun\n          as\u00ed prefieres no nombrar lo que no hab\u00e9is entregado, borra\n          esas celdas y deja el guion largo. >>\n       ============================================================ -->\n  <section class=\"tb-sec tb-ins-group\" id=\"cloud-map\" data-n=\"07\">\n    <div class=\"tb-ins-col\">\n      <p class=\"tb-lead-label\">07 \u00b7 Cross-cloud map<\/p>\n      <h2 class=\"tb-h2\">The same job, <em>three different names.<\/em><\/h2>\n      <p class=\"tb-lede\">\n        Most of the confusion in a data conversation comes from one thing: each cloud sells\n        the same capability under a different brand. This is the translation table.\n      <\/p>\n    <\/div>\n\n    <div class=\"tb-ins-map-wrap\">\n      <table class=\"tb-ins-map\">\n        <caption>\n          These are rough equivalents, not identical products. They are close enough to follow\n          a commercial conversation and not close enough to migrate between without thinking.\n        <\/caption>\n        <thead>\n          <tr>\n            <th scope=\"col\">The job<\/th>\n            <th scope=\"col\">Amazon Web Services<\/th>\n            <th scope=\"col\">Microsoft Azure<\/th>\n            <th scope=\"col\">Google Cloud<\/th>\n            <th scope=\"col\">Cloud-independent<\/th>\n          <\/tr>\n        <\/thead>\n        <tbody>\n          <tr>\n            <th scope=\"row\">Object storage<\/th>\n            <td>Amazon S3<\/td>\n            <td>Blob Storage \/ ADLS<\/td>\n            <td>Cloud Storage (GCS)<\/td>\n            <td class=\"tb-ins-none\">\u2014<\/td>\n          <\/tr>\n          <tr>\n            <th scope=\"row\">Warehouse &amp; analytics<\/th>\n            <td>Amazon Redshift<\/td>\n            <td>Synapse \/ Fabric<\/td>\n            <td>BigQuery<\/td>\n            <td class=\"tb-ins-multi\">Snowflake<\/td>\n          <\/tr>\n          <tr>\n            <th scope=\"row\">Large-scale processing<\/th>\n            <td>Databricks \/ EMR<\/td>\n            <td>Azure Databricks<\/td>\n            <td>Databricks \/ Dataproc<\/td>\n            <td class=\"tb-ins-multi\">Databricks + Spark<\/td>\n          <\/tr>\n          <tr>\n            <th scope=\"row\">Orchestration<\/th>\n            <td>Amazon MWAA<\/td>\n            <td>Azure Data Factory<\/td>\n            <td>Cloud Composer<\/td>\n            <td class=\"tb-ins-multi\">Apache Airflow<\/td>\n          <\/tr>\n          <tr>\n            <th scope=\"row\">Streaming &amp; events<\/th>\n            <td>Kinesis \/ MSK<\/td>\n            <td>Event Hubs<\/td>\n            <td>Pub\/Sub<\/td>\n            <td class=\"tb-ins-multi\">Apache Kafka<\/td>\n          <\/tr>\n          <tr>\n            <th scope=\"row\">SQL transformation<\/th>\n            <td>dbt<\/td>\n            <td>dbt<\/td>\n            <td>dbt<\/td>\n            <td class=\"tb-ins-multi\">dbt<\/td>\n          <\/tr>\n        <\/tbody>\n      <\/table>\n    <\/div>\n\n    <div class=\"tb-ins-col\">\n      <p class=\"tb-ins-up\"><a href=\"#index\">\u2191 Back to the index<\/a><\/p>\n    <\/div>\n  <\/section>\n\n\n  <!-- ============================================================\n       09 \u00b7 CONOCIMIENTO TRANSFERIBLE\n       Esta secci\u00f3n es puro argumento comercial disfrazado de\n       contenido \u00fatil \u2014 y es honesto, porque es verdad. Responde por\n       adelantado a la objeci\u00f3n \"\u00bfpero hab\u00e9is trabajado con ESTA\n       herramienta exacta?\", que es la que m\u00e1s proyectos os va a\n       costar en los pr\u00f3ximos dos a\u00f1os.\n       ============================================================ -->\n  <section class=\"tb-sec tb-ins-group\" id=\"transferable\" data-n=\"08\">\n    <div class=\"tb-ins-col\">\n      <p class=\"tb-lead-label\">08 \u00b7 A note on experience<\/p>\n      <h2 class=\"tb-h2\">Most of this knowledge <em>transfers.<\/em><\/h2>\n      <p class=\"tb-lede\">\n        When a requirement names a specific tool, the useful question is rarely\n        \u201chas this person used exactly this product?\u201d It is\n        \u201cdo they understand the concept underneath it?\u201d\n      <\/p>\n\n      <article class=\"tb-ins-term\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Airflow \u2192 Cloud Composer<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Near-identical<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">\n          Composer <em>is<\/em> Airflow, managed by Google. Someone fluent in Airflow is\n          productive in Composer almost immediately.\n        <\/p>\n      <\/article>\n\n      <article class=\"tb-ins-term\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">Amazon S3 \u2192 Google Cloud Storage<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Same concept<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">\n          The idea of object storage carries over intact. What has to be learned is the\n          surrounding Google Cloud specifics, not the model itself.\n        <\/p>\n      <\/article>\n\n      <article class=\"tb-ins-term\">\n        <div class=\"tb-ins-head\">\n          <h3 class=\"tb-ins-t\">BigQuery \u2192 Snowflake<\/h3>\n          <span class=\"tb-ins-lead\" aria-hidden=\"true\"><\/span>\n          <span class=\"tb-ins-cat\">Related, not equal<\/span>\n        <\/div>\n        <p class=\"tb-ins-d\">\n          Many concepts overlap and transfer well. They are still different products with\n          different cost models and different operational behaviour, and anyone presenting\n          them as identical is skipping the part that matters.\n        <\/p>\n      <\/article>\n\n      <p class=\"tb-ins-d\" style=\"margin-top:26px\">\n        This cuts both ways, and it is why we say it here rather than in a sales call: it\n        means a stack mismatch is often not a real problem, and it also means\n        \u201cwe know that tool\u201d is a weaker claim than it sounds. Ask what the\n        concept is, not what the logo is.\n      <\/p>\n\n      <p class=\"tb-ins-up\"><a href=\"#index\">\u2191 Back to the index<\/a><\/p>\n    <\/div>\n  <\/section>\n\n\n  <!-- ============================================================\n       10 \u00b7 COLOF\u00d3N\n       Cierre propio de esta p\u00e1gina: la ficha t\u00e9cnica del documento.\n       La fecha de revisi\u00f3n no es decorativa \u2014 una obra de referencia\n       con fecha visible es m\u00e1s citable que una sin ella.\n       << Rellena la fecha con el d\u00eda que lo publiques. Y actual\u00edzala\n          de verdad cada vez que a\u00f1adas o cambies un t\u00e9rmino. >>\n       ============================================================ -->\n  <section class=\"tb-ins-end\">\n    <div class=\"tb-ins-end-in\">\n      <div class=\"tb-ins-end-cell\">\n        <p class=\"tb-ins-end-k\">Last reviewed<\/p>\n        <p class=\"tb-ins-end-v\">&lt;&lt;2026-08-26&gt;&gt;<\/p>\n      <\/div>\n      <div class=\"tb-ins-end-cell\">\n        <p class=\"tb-ins-end-k\">Written and maintained by<\/p>\n        <p class=\"tb-ins-end-v\">Enrique Delgado Aznar and Juan Navarro Micol, from real project conversations.<\/p>\n      <\/div>\n      <div class=\"tb-ins-end-cell\">\n        <p class=\"tb-ins-end-k\">Missing a term?<\/p>\n        <p class=\"tb-ins-end-v\"><a href=\"\/contact\/\">Tell us and we&#8217;ll add it <i>\u2192<\/i><\/a><\/p>\n      <\/div>\n    <\/div>\n  <\/section>\n\n<\/div>\n\n\n<!-- =====================================================================\n     JSON-LD \u00b7 DefinedTermSet con los 29 t\u00e9rminos\n     ---------------------------------------------------------------------\n     DefinedTermSet + DefinedTerm es el tipo que schema.org tiene\n     EXACTAMENTE para esto: un glosario y sus entradas. Cada t\u00e9rmino va\n     con su @id apuntando a su ancla, as\u00ed que un motor puede citar\n     tobolist.com\/insights\/#apache-kafka y no solo la p\u00e1gina entera.\n\n     AVISO PARA CUANDO LO VALIDES\n     La Prueba de resultados enriquecidos probablemente dir\u00e1 \"no se ha\n     detectado ning\u00fan resultado enriquecido\". ESO NO ES UN ERROR.\n     DefinedTermSet no genera tarjeta enriquecida en el buscador: es un\n     tipo de datos, no de presentaci\u00f3n. Lo que importa es que salga\n     0 errores y 0 avisos. Si quieres verlo interpretado, usa el\n     Validador de marcado de schema.org (validator.schema.org), que s\u00ed\n     lista los 29 t\u00e9rminos.\n\n     Es la lecci\u00f3n del fichero 27 aplicada por adelantado: elegimos el\n     tipo por lo que significa, no por si sale una tarjeta bonita.\n\n     << Rellena \"dateModified\" con la fecha de publicaci\u00f3n, la misma\n        que pongas en el colof\u00f3n. >>\n     ===================================================================== -->\n\n<script type=\"application\/ld+json\">\n{\n  \"@context\": \"https:\/\/schema.org\",\n  \"@graph\": [\n    {\n      \"@type\": \"Organization\",\n      \"@id\": \"https:\/\/tobolist.com\/#organization\",\n      \"name\": \"Tobolist\",\n      \"url\": \"https:\/\/tobolist.com\/\"\n    },\n    {\n      \"@type\": \"DefinedTermSet\",\n      \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\",\n      \"name\": \"The Tobolist data glossary\",\n      \"description\": \"A plain-language glossary of the roles, platforms and concepts that come up in data engineering, data architecture and business intelligence projects \u2014 written for the people who have to buy, approve or explain them.\",\n      \"url\": \"https:\/\/tobolist.com\/insights\/\",\n      \"inLanguage\": \"en\",\n      \"publisher\": {\n        \"@id\": \"https:\/\/tobolist.com\/#organization\"\n      },\n      \"author\": {\n        \"@id\": \"https:\/\/tobolist.com\/#organization\"\n      },\n      \"dateModified\": \"<<AAAA-MM-DD>>\",\n      \"hasDefinedTerm\": [\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#data-architect\",\n          \"name\": \"Data architect\",\n          \"description\": \"Decides how the data platform should be built, before anyone builds it. Chooses the technologies and architecture patterns, defines how data is stored, processed and integrated, and sets the standards on scalability, security and governance that the engineering team then follows.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#data-architect\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#data-engineer\",\n          \"name\": \"Data engineer\",\n          \"description\": \"Builds and maintains the infrastructure that moves, transforms, stores and processes data. The job is to get data out of the systems where it lives \u2014 ERP, CRM, internal apps, databases \u2014 and into the place where it will actually be used, reliably and on schedule. Typical stack: Python, SQL, Spark, Databricks, Kafka, Airflow, dbt, and one of the three big clouds.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#data-engineer\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#bi-analyst\",\n          \"name\": \"BI analyst\",\n          \"description\": \"Turns available data into something a business can decide with: KPIs, dashboards, reports, reporting-oriented data models. Sits closest to the business user of the three roles \u2014 the job is not to build the platform but to make the numbers legible and defensible. Typical stack: Power BI, Tableau, Looker, SQL, Excel, Microsoft Fabric.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#bi-analyst\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#databricks\",\n          \"name\": \"Databricks\",\n          \"description\": \"A platform for processing, engineering and analysing large volumes of data, built around Apache Spark. Different profiles \u2014 engineers, analysts, data scientists \u2014 work on the same platform to transform data, run jobs, train machine learning models and build data products. It is not tied to one cloud: it runs on Azure, AWS and Google Cloud alike.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#databricks\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#apache-spark\",\n          \"name\": \"Apache Spark\",\n          \"description\": \"The open-source engine for distributed data processing that sits underneath a large part of the modern data stack, Databricks included. When a job is too big for one machine, Spark splits it across many. PySpark is its Python interface \u2014 the one most data engineers actually write.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#apache-spark\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#snowflake\",\n          \"name\": \"Snowflake\",\n          \"description\": \"A cloud platform specialised in storing, processing and querying large amounts of data for analytics, reporting, BI and data science. Like Databricks, it is independent of any single cloud: it runs on top of AWS, Azure or Google Cloud infrastructure rather than belonging to any of them.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#snowflake\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#google-bigquery\",\n          \"name\": \"Google BigQuery\",\n          \"description\": \"Google Cloud's managed data warehouse and analytics platform. Stores very large volumes of data and runs queries over them without you managing servers. If a company has chosen Google Cloud, BigQuery is usually the centre of its analytics.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#google-bigquery\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#azure-synapse\",\n          \"name\": \"Azure Synapse Analytics\",\n          \"description\": \"Microsoft Azure's analytics and data processing platform, combining storage, processing and analysis. For years it was the centre of Microsoft's data ecosystem. Microsoft is now pushing Microsoft Fabric, so new Azure projects increasingly mention Fabric instead.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#azure-synapse\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#microsoft-fabric\",\n          \"name\": \"Microsoft Fabric\",\n          \"description\": \"Microsoft's unified analytics platform, bringing storage, data engineering, warehousing and Power BI reporting into one product. It is where Microsoft is steering new Azure data projects, and it is increasingly the default answer inside Microsoft-heavy organisations.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#microsoft-fabric\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#amazon-redshift\",\n          \"name\": \"Amazon Redshift\",\n          \"description\": \"AWS's managed data warehouse. Plays the role that BigQuery plays on Google Cloud and that Synapse or Fabric play on Azure: the place where structured data is stored and queried for analytics.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#amazon-redshift\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#data-lake\",\n          \"name\": \"Data lake\",\n          \"description\": \"A repository that holds data as it arrives, in whatever shape it arrives: files, logs, exports, images, raw dumps. Cheap and flexible, because nothing has to be modelled before it lands. The cost is that a lake with no governance quietly becomes a place where data goes to be forgotten.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#data-lake\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#data-warehouse\",\n          \"name\": \"Data warehouse\",\n          \"description\": \"A repository that holds structured, modelled data ready for analysis: sales by customer, monthly billing, margins by region. Data has to be cleaned and shaped before it goes in, which is more work up front and far less work every time somebody asks a question.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#data-warehouse\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#lakehouse\",\n          \"name\": \"Lakehouse\",\n          \"description\": \"An architecture that tries to keep the cheap, flexible storage of a lake and add the structure, reliability and query performance of a warehouse on top of it. It is the pattern behind most modern platform designs, and the one Databricks built its product around.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#lakehouse\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#amazon-s3\",\n          \"name\": \"Amazon S3\",\n          \"description\": \"AWS's object storage service, and one of the foundations of most data architectures built on AWS: files, datasets, logs, backups, raw data, processing output. Very often it is the data lake.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#amazon-s3\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#google-cloud-storage\",\n          \"name\": \"Google Cloud Storage\",\n          \"description\": \"Google Cloud's object storage service, usually shortened to GCS. Does essentially the same job as Amazon S3 within the Google Cloud ecosystem.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#google-cloud-storage\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#azure-blob-storage\",\n          \"name\": \"Azure Blob Storage \/ ADLS\",\n          \"description\": \"Microsoft Azure's object storage. ADLS \u2014 Azure Data Lake Storage \u2014 is the variant built for analytics workloads. Together they are the Azure equivalent of S3 or GCS.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#azure-blob-storage\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#data-pipeline\",\n          \"name\": \"Data pipeline\",\n          \"description\": \"The set of automated steps that take data from where it is produced to where it is consumed: extract it, clean it, transform it, load it, and do it again tomorrow without anybody pressing a button. When people say a data project failed, they usually mean the pipeline stopped being trustworthy.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#data-pipeline\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#etl-elt\",\n          \"name\": \"ETL \/ ELT\",\n          \"description\": \"Two orderings of the same three steps. ETL \u2014 extract, transform, load \u2014 cleans the data before storing it. ELT \u2014 extract, load, transform \u2014 stores it raw first and transforms it inside the destination platform, which is what modern cloud warehouses are fast enough to allow.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#etl-elt\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#apache-airflow\",\n          \"name\": \"Apache Airflow\",\n          \"description\": \"The tool that organises, schedules and supervises data processes. Think of a conductor: at 02:00 pull this data, when it finishes run this transformation, then refresh this table, and raise an alert if any step fails. That job is called pipeline orchestration.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#apache-airflow\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#cloud-composer\",\n          \"name\": \"Google Cloud Composer\",\n          \"description\": \"Google Cloud's managed Airflow service. Composer is Airflow, run and maintained by Google.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#cloud-composer\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#azure-data-factory\",\n          \"name\": \"Azure Data Factory\",\n          \"description\": \"Microsoft Azure's service for integrating, moving and orchestrating data between systems \u2014 ERP, SQL Server, files, APIs, external applications \u2014 and delivering it into the data platform. Usually shortened to ADF, and near-universal in Azure-based companies.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#azure-data-factory\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#dbt\",\n          \"name\": \"dbt\",\n          \"description\": \"A tool for transforming and organising data using mostly SQL. The data has usually already landed in a warehouse, and dbt turns it from raw tables into structured, documented, tested business concepts: sales by customer, monthly billing, margin by region.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#dbt\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#batch-vs-streaming\",\n          \"name\": \"Batch vs streaming\",\n          \"description\": \"Batch means processing data in scheduled blocks: every night we process the day's sales. Streaming means processing events as they happen: we want to see each sale as it occurs. Streaming is more expensive to build and to run, and it is genuinely necessary far less often than it gets asked for.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#batch-vs-streaming\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#apache-kafka\",\n          \"name\": \"Apache Kafka\",\n          \"description\": \"The best-known technology for carrying large volumes of information continuously and in near real time: transactions, application events, sensor readings, user activity, logs. If Kafka is in the requirements, the project is usually a fairly technical data engineering job.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#apache-kafka\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#pubsub\",\n          \"name\": \"Google Cloud Pub\/Sub\",\n          \"description\": \"Google Cloud's managed service for sending and receiving events or messages between systems, used in real-time and event-driven architectures.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#pubsub\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#power-bi\",\n          \"name\": \"Power BI\",\n          \"description\": \"Microsoft's business intelligence and data visualisation tool, and by some distance the most common reporting layer in European mid-market companies. Connects to data sources, models the data and produces the dashboards a management team reads.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#power-bi\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#kpi\",\n          \"name\": \"KPI\",\n          \"description\": \"A key performance indicator: a single number that is supposed to tell you whether something is going well. The hard part is almost never calculating it \u2014 it is agreeing on its definition across departments so that finance and sales don't produce two different revenue figures from the same data.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#kpi\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#semantic-model\",\n          \"name\": \"Semantic model\",\n          \"description\": \"The layer that defines what each business concept means in the reporting tool: what counts as revenue, what counts as an active customer, how a margin is calculated. It is where a definition lives once instead of being re-invented in every dashboard.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#semantic-model\"\n        },\n        {\n          \"@type\": \"DefinedTerm\",\n          \"@id\": \"https:\/\/tobolist.com\/insights\/#data-quality\",\n          \"name\": \"Data quality\",\n          \"description\": \"Whether the data is complete, correct, consistent and current enough to be used. Quality problems that are only detected in the reporting layer have already cost you the decision that was made on them, which is why checks belong upstream in the pipeline.\",\n          \"inDefinedTermSet\": {\n            \"@id\": \"https:\/\/tobolist.com\/insights\/#glossary\"\n          },\n          \"url\": \"https:\/\/tobolist.com\/insights\/#data-quality\"\n        }\n      ]\n    }\n  ]\n}\n<\/script>\n","protected":false},"excerpt":{"rendered":"<p>Services Work Insights About Contact Insights The data glossary, written for the person paying for it. Every data proposal you receive will be full of names: Databricks, Snowflake, Airflow, Fabric, lakehouse, streaming. Most glossaries explain them to engineers. This one explains them to whoever has to decide, approve or defend the project \u2014 including which &#8230; <a title=\"Perceptions\" class=\"read-more\" href=\"https:\/\/tobolist.com\/es\/insights\/\" aria-label=\"Leer m\u00e1s sobre Insights\">Leer m\u00e1s<\/a><\/p>","protected":false},"author":1,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"footnotes":""},"class_list":["post-272","page","type-page","status-publish"],"_links":{"self":[{"href":"https:\/\/tobolist.com\/es\/wp-json\/wp\/v2\/pages\/272","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/tobolist.com\/es\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/tobolist.com\/es\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/tobolist.com\/es\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/tobolist.com\/es\/wp-json\/wp\/v2\/comments?post=272"}],"version-history":[{"count":2,"href":"https:\/\/tobolist.com\/es\/wp-json\/wp\/v2\/pages\/272\/revisions"}],"predecessor-version":[{"id":407,"href":"https:\/\/tobolist.com\/es\/wp-json\/wp\/v2\/pages\/272\/revisions\/407"}],"wp:attachment":[{"href":"https:\/\/tobolist.com\/es\/wp-json\/wp\/v2\/media?parent=272"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}