Data Engineer
We are seeking a seasoned Data Engineer who brings deep, hands-on expertise in SQL and relational database systems. The ideal candidate developed their craft over at least 6 years, possessing battle-tested fundamentals in query design, schema architecture, and performance tuning. While the role centers on SQL and data infrastructure, we expect a full-stack awareness; someone who understands how data flows from ingestion through transformation to consumption and can collaborate credibly across the entire engineering stack. You will help architect an AI-native platform that includes knowledge graph construction, LLM-powered entity resolution, and agent-driven query layers.
- Education
- 4 year undergraduate degree
- Experience
- 6+ years
- Location
- Hoboken, NJ metro area
- Work model
- On-site, 5 days per week
- Classification
- Full-stack mindset; SQL / data specialization
Technical core
- Advanced SQL mastery — Window functions, CTEs (recursive and non-recursive), subquery optimization, set-based logic, complex joins, and dynamic SQL.
- Query performance engineering — Execution plan analysis (EXPLAIN / EXPLAIN ANALYZE), indexing, partitioning, query rewrites, and statistics-driven tuning.
- Database administration — Backup and restore strategies, replication, connection pooling, vacuum tuning, role-based access control, and disaster recovery planning.
- ETL/ELT pipeline development — Hands-on experience designing, building, and maintaining production-grade data pipelines using modern orchestration and transformation tooling.
Database technologies
- PostgreSQL — Strongly preferred. PL/pgSQL stored procedures, triggers, materialized views, JSONB querying, full-text search, extensions ecosystem.
- Additional RDBMS — Microsoft SQL Server (T-SQL, SSMS, SSIS, SSRS, SSAS), MySQL/MariaDB, or Oracle PL/SQL experience is a plus.
- Cloud data platforms — Familiarity with Snowflake, Google BigQuery, Amazon Redshift, or Azure Synapse Analytics.
Tooling and infrastructure
- Orchestration — Apache Airflow or Prefect for DAG-based workflow orchestration and scheduling.
- Python — pandas, SQLAlchemy, psycopg2/asyncpg, Alembic for migrations, Great Expectations for data quality.
- Scripting — Bash and shell scripting for automation and cron-based workflows.
- DevOps — Git version control (GitHub/GitLab), CI/CD for data pipelines (Azure DevOps), Docker and containerized database environments.
AI, graph and knowledge engineering
- Graph databases — Experience or strong familiarity with Neo4j, Apache AGE, or similar. Proficiency in Cypher or Gremlin, and understanding of graph modeling, entity resolution, and relationship traversal.
- LLM integration — Practical exposure to integrating large language models into data pipelines: prompt engineering, retrieval-augmented generation, and orchestrating LLM-powered entity extraction and classification.
- Vector and embedding infrastructure — Familiarity with vector databases (pgvector, Pinecone, Weaviate) and embedding-based search for semantic retrieval and similarity matching.
- Agent frameworks — Awareness of agentic frameworks (LangChain, LlamaIndex, Semantic Kernel) and how they connect to underlying data infrastructure is a strong plus.
Visualization and business intelligence
- Tableau — Nice to have. Connecting live and extract data sources, calculated fields, LOD expressions, dashboard design, Tableau Server/Cloud publishing, and performance optimization.
Attributes
- Years of building complex queries and data models
- Comfortable working cross-functionally with front-end engineers and analysts
- Clear communicator who can translate complex data concepts for non-technical stakeholders
- Interest in emerging data technologies: real-time streaming (Kafka, Flink), vector databases, or data mesh architectures