Apache Hop Visual Data Orchestration

AWS Analytics

Apache Hop, the Hop Orchestration Platform, for visually designing and running data pipelines and workflows, with per instance credentials and a dedicated data volume, backed by cloudimg support.

Base
Hardened build
minimal ports, security patches applied at build time
Access
Unique credentials
generated on first boot, readable only by root
Verified
Boots working
services pass a health gate before release
Support
24/7, 365 days
by email and live chat, 24 hour response SLA

Overview

Apache Hop, the Hop Orchestration Platform, is an open source data engineering and data orchestration platform for visually designing and running data pipelines and workflows. It provides a browser based, drag and drop editor for building, running and monitoring data integration, transformation and orchestration logic, plus a headless execution server for scheduled and remote pipeline runs. This image ships Hop fully installed and configured so a complete data orchestration platform is ready within minutes of launch: the Hop Web visual editor behind an authenticated reverse proxy, and the Hop Server execution server, both supervised by the operating system.

Why the cloudimg image

cloudimg generates a fresh password on first boot for every instance and applies it to both the web editor and the execution server, replacing the well known stock defaults so no two instances ship the same credentials. Hop projects, pipelines, workflows and the execution server configuration live on a dedicated data volume that can be resized independently of the operating system disk. Every image is paired with a step by step deploy guide and is backed by 24/7 cloudimg support.

Common uses

  • Data integration and ETL between databases, APIs, files and cloud storage
  • Scheduled and orchestrated multi step data workflows run on the execution server
  • Migrating and modernising existing Pentaho and Kettle pipelines

Key features

  • Unlike a manual install that takes hours of Java, Docker and systemd configuration, this AMI delivers Apache Hop fully running in minutes - the Hop Web visual editor and Hop Server execution engine are pre-wired behind authenticated endpoints on a dedicated EBS data volume, so you start building pipelines immediately with zero setup.
  • Unlike stock Hop images that ship well-known default credentials exposed to the internet, this AMI generates a unique password on every first boot and applies it to both the web editor and execution server automatically - credentials are stored in a root-only file, closing the most common security gap in self-managed deployments.
  • Unlike self-supported open source installs, this AMI includes 24/7 expert support from cloudimg with a one-hour average response for critical issues - our engineers help with Hop deployment, pipeline design, execution server configuration, Pentaho/Kettle migration and performance tuning.

See it running

Real screenshots taken while testing this image against its deployment guide.

Apache Hop Visual Data Orchestration screenshot 1 Apache Hop Visual Data Orchestration screenshot 2 Apache Hop Visual Data Orchestration screenshot 3 Apache Hop Visual Data Orchestration screenshot 4

Description

This is a repackaged open source software product wherein additional charges apply for cloudimg support services.

Overview Apache Hop (the Hop Orchestration Platform) is an open source data engineering and data orchestration platform for visually designing and running data pipelines and workflows. It provides a browser-based, drag-and-drop editor to build, run and monitor data integration, transformation and orchestration logic, plus a headless execution server for scheduled and remote pipeline runs. This AMI delivers Hop fully installed and configured so a complete data orchestration platform is running within minutes of launch. The current release available is Apache Hop. About cloudimg cloudimg specialises in delivering production-hardened, support-backed AMIs on AWS Marketplace. Our engineering team maintains and supports every image we publish, ensuring rapid updates, security patching and hands-on guidance for customers running data workloads in production. Why This AMI Over Alternatives Unlike a bare Hop installation, a generic community image, or a Docker-based setup, this AMI eliminates the manual steps that slow down production readiness: - No shared default credentials - A one-shot first-boot service generates a unique password per instance and writes it into both the web editor and the execution server, replacing the well-known stock defaults that ship with a manual install. This closes the security risk of default credentials being exposed on the public internet.

  • Two runtimes, pre-wired - The Hop Web visual editor and the Hop Server execution server are both installed, ordered and supervised by systemd, so you can design a pipeline in the browser and run it on the server without wiring anything together yourself.
  • Independently resizable storage - Hop projects, pipelines, workflows and the execution server configuration live on a dedicated EBS data volume, so you can resize storage without rebuilding the instance.
  • Immediate productivity - Launch, retrieve your credentials, open the editor and start building pipelines without touching Java configuration, Docker, systemd units or reverse-proxy setup. Application Stack Apache Hop is installed from the official Apache distribution. The Hop Web visual editor runs as the official Apache Hop web container, reverse-proxied by nginx on port 80 with HTTP Basic authentication. The Hop Server execution server runs under a dedicated unprivileged service account and is reachable only through its own authenticated endpoint. Both runtimes bind the loopback interface and are fronted with authentication; the Hop projects, pipelines, workflows and server configuration live on a dedicated, independently resizable EBS data volume. Runs on OpenJDK 21. Secure First Boot On first boot a one-shot service generates a fresh password unique to that instance, applies it to both the web editor and the execution server, and writes the credentials to a root-only file. No shared or default credentials ship in the image. Because the Hop execution server can run pipelines remotely, replacing its stock default credentials on every instance is essential, and this AMI does it automatically. Getting Started 1. Launch the AMI on your chosen EC2 instance type.

2. SSH into the instance and retrieve your generated credentials from the root-only file.

3. Browse to the instance public address and sign in to the Hop Web editor.

4. Open the bundled sample project, then design and run your first pipeline. For evaluation, launch on a smaller instance type to explore the editor and test pipeline logic at minimal cost before scaling up for production workloads. Ready to get started with expert guidance? Contact cloudimg for a free deployment consultation - our engineers can walk you through your first Hop pipeline, help plan a migration from Pentaho/Kettle, or advise on instance sizing for your workload. Example Use Case: Batch ETL Into a Data Warehouse A data team can use Hop to read from operational databases and flat files, apply cleansing, lookups and transformations visually, and load the results into a warehouse or object storage on a schedule. The Hop Server runs the workflow unattended and exposes execution status, so engineers monitor and adjust pipelines without redeploying code. Additional Use Cases - Data integration between databases, APIs, files and cloud storage

  • Scheduled and orchestrated multi-step data workflows
  • Migration and consolidation of Pentaho/Kettle pipelines
  • Reusable, parameterised pipelines promoted across environments cloudimg Support 24/7 technical support by email and live chat. Our engineers assist with Hop deployment, upgrades, pipeline and workflow design, execution server configuration and performance tuning. Critical issues receive a one-hour average response time. Apache Hop, Hop and Apache are trademarks of the Apache Software Foundation.

Related technologies

data orchestration platformvisual pipeline editoretl amidata integration amihop serverpentaho kettle migrationdata engineeringbatch etlworkflow automationdata pipeline aws