We provide IT Staff Augmentation Services!

Senior Big Data Engineer Resume

0/5 (Submit Your Rating)

West Chester, PA

SUMMARY

  • Having 8+ Years of strong experience in Software Development Life Cycle (SDLC) including Requirements Analysis, Design Specification and Testing as per Cycle in both Waterfall and Agile methodologies.
  • Good exposure working with Hadoop distributions such Cloudera, Hortonworks and, Data Bricks.
  • Hands on experience in various languages like Java, Scala, Python and Linux/Unix shell scripting
  • Experience in developing customizedUDF’sin Python to extend Hive and Pig Latin functionality.
  • Proficient in data mining tools like R, SAS, Python, SQL, Excel, Big Data Hadoop eco - systems Staff leadership and development
  • Strong experience in Big Data Analytics using HDFS, YARN, MapReduce, Hive, Impala, Pig, Sqoop, HBase, Spark, Spark SQL, Kafka, Spark Streaming, Flume, Oozie, Zookeeper, Hue.
  • Extensive noledge in writing Hadoop jobs for data analysis as per teh business requirements using Hive and worked on HiveQL queries for required data extraction, join operations, writing custom UDF's as required and having good experience in optimizing Hive Queries.
  • Experience in PySpark programming language with Spark Core and Spark modules extensively.
  • Experience in dealing with data formats ORC, Parquet, JSON and CSV.
  • Experience in performance tuning by using Partitioning, Bucketing and Indexing in Hive.
  • Experienced in job workflow scheduling and monitoring tools like Airflow, Oozie, TWS, Control-M and Zookeeper.
  • Developed custom Kafka producer and consumer for different publishing and subscribing to Kafka topics.
  • Good working experience on Spark (spark streaming, spark SQL) with Scala and Kafka. Worked on reading multiple data formats on HDFS using Scala.
  • Good working noledge of Amazon Web Services(AWS) Cloud Platform which includes services likeEC2,S3,VPC,ELB, IAM, DynamoDB, Cloud Front, Cloud Watch, Route 53, Elastic Beanstalk (EBS), Auto Scaling, Security Groups.
  • Implemented various frameworks for data pipelines and workflows using HBase, Kafka with Spark/PySpark, Python and Scala.
  • Extensively worked on Spark using Scala on cluster for computational (analytics), installed it on top of Hadoop performed advanced analytical application by making use of Spark with Hive and SQL/Oracle/Snowflake
  • Experience in Microsoft Azure/Cloud Services like SQL Data Warehouse, Azure SQL Server, Azure Databricks, Azure Data Lake, Azure Blob Storage, Azure Data Factory.
  • Good noledge in OLAP, OLTP, Business Intelligence and Data Warehousing concepts with emphasis on ETL and Business Reporting needs.
  • Extensive experience working on various databases and database script development using SQL and PL/SQL.
  • Develop generic SQL Procedures and Complex T-SQL statements to achieve teh reports generation.
  • Perform structural modifications using Map - Reduce, HIVE and analyze data using visualization/ reporting tools.
  • Experience in developing customized UDF's in Python to extend Hive and Pig Latin functionality.
  • Excellent noledge of Microsoft Office with an emphasis on Excel.
  • Experience in Performance Tuning and query optimization techniques in transactional and Data Warehouse Environments.
  • Experience with Software development tools such as JIRA, GIT, SVN.
  • Working experience with Linux lineup like Redhat and CentOS.
  • Experience on ETL concepts using Informatica Power Center, AB Initio.

TECHNICAL SKILLS

Hadoop Distribution: Cloudera CDH, Apache, AWS, Horton Works HDP

Big Data Tools: Kafka, Cassandra, Apache Spark, Spark Streaming, HBase, Impala, HDFS, MapReduce, Hive, Pig, Sqoop, Flume, Oozie, Zookeeper

Programming Languages: SQL, PL/SQL, Python, UNIX, Pyspark, Pig, HiveQL, Scala, Shell Scripting

Spark Components: RDD, Spark SQL, Spark Streaming

Data Modeling Tools: Erwin Data Modeler, ER Studio v17

Methodologies: RAD, JAD, System Development Life Cycle (SDLC), Agile

Cloud Management: MS Azure, Amazon Web Services (AWS)- EC2, EMR, S3, Redshift, EMR, Lambda, Atana

Databases: Oracle 12c/11g/ 10g, MySql, MS Sql, DB2, Snowflake

No Sql Databases: MongoDB, Hbase, Cassandra

OLAP Tools: Tableau, SSAS, Business Objects, and Crystal Reports 9

ETL/Data warehouse Tools: Informatica, and Tableau.

Version Control: CVS, SVN, Clear Case, Git

Operating System: Windows, Unix, Sun Solaris

PROFESSIONAL EXPERIENCE:

Confidential, West Chester, PA

Senior Big Data Engineer

Responsibilities:

  • Used Airflow for scheduling teh Hive, Spark and MapReduce jobs.
  • Developed Spark/Scala, Python for regular expression (regex) project in teh Hadoop/Hive environment with Linux/Windows for big data resources.
  • Data sources are extracted, transformed and loaded to generate CSV data files with Python programming and SQL queries.
  • Subscribing teh Kafka topic with Kafka consumer client and process teh events in real time using spark.
  • Collected data using Spark Streaming from AWS S3 bucket in near-real-time and performs necessary Transformations and Aggregation on teh fly to build teh common learner data model and persists teh data in HDFS.
  • Designed and Developed Real Time Stream Processing Application using Spark, Kafka, Scala and Hive to perform Streaming ETL and apply Machine Learning.
  • Use SparkSQL to load JSON data and create Schema RDD and loaded it into Hive Tables and handled structured data using SparkSQL.
  • Developed Spark code using Scala and Spark-SQL/Streaming for faster processing of data.
  • Developing Spark programs with Python, and applied principals of functional programming to process teh complex structured data sets.
  • Worked with Hadoop infrastructure to storage data in HDFS storage and use Spark / HIVE SQL to migrate underlying SQL codebase in AWS.
  • Developed reusable objects like PL/SQL program units and libraries, database procedures and functions, database triggers to be used by teh team and satisfying teh business rules.
  • Involved with writing scripts in Oracle, SQL Server and Netezza databases to extract data for reporting and analysis and Worked in importing and cleansing of data from various sources like DB2, Oracle, flat files onto SQL Server with high volume data
  • Worked with Hadoop ecosystem and Implemented Spark using Scala and utilized Data frames and Spark SQL API for faster processing of data.
  • Using Python in spark to extract teh data from Snowflake and upload it to Salesforce on Daily basis.
  • Developed Spark Streaming job to consume teh data from teh Kafka topic of different source systems and push teh data into HDFS locations.
  • Designed and developed architecture for data services ecosystem spanning Relational, NoSQL, and Big data technologies. Extracted Mega Data from Amazon Redshift, AWS, and Elastic Search engine using SQL Queries to create reports.
  • Used Talend for Big Data Integration using Spark and Hadoop.
  • Responsible for analyzing large data sets and derive customer usage patterns by developing new MapReduce programs using Java.
  • Analyzing SQL scripts and designed teh solution to implement using PySpark
  • Export tables from Teradata to HDFS using Sqoop and build tables in Hive.
  • Loaded and transformed large sets of structured, semi structured and unstructured data usingHadoop/Big Data concepts.
  • Converting Hive/SQL queries into Spark transformations using Spark RDDs and Pyspark
  • Filtering and cleaning data using Scala code and SQL Queries
  • Troubleshooting errors in Hbase Shell/API, Pig, Hive and MapReduce.
  • Implemented Installation and configuration of multi-node cluster on Cloud using Amazon Web Services (AWS) onEC2.
  • Use python to write a service which is event based using AWS Lambda to achieve real time data to One-Lake (A Data Lake solution in Cap-One Enterprise).
  • Developed Automation Regressing Scripts for validation of ETL process between multiple databases like AWS Redshift, Oracle, MongoDB, T-SQL, and SQL Server usingPython.
  • Designed Kafka producer client using Confluent Kafka and produced events into Kafka topic.
  • Worked on SQL Server concepts SSIS (SQL Server Integration Services), SSAS (Analysis Services) and SSRS (Reporting Services). Using Informatica & SSIS, SPSS, SAS to extract transform & load source data from transaction systems.

Environment: Hadoop, Spark, Scala, Hbase, Hive, Python, Snowflake, PL/SQL AWS, EC2, S3, Lambda, Auto Scaling, Cloud Watch, Cloud Formation, Confidential Info sphere, DataStage, MapReduce, Oracle12c, Flat files, TOAD, MS SQL Server database, XML files, Cassandra, MongoDB, Kafka, MS Access database, Autosys, UNIX, Erwin.

Confidential, Urbandale, IA

Big Data Engineer

Responsibilities:

  • Worked on reading and writing multiple data formats like JSON, ORC, Parquet on HDFS using PySpark.
  • Provide guidance to development team working on PySpark as ETL platform
  • Used Azure Databricks for fast, easy, and collaborative spark-based platform on Azure.
  • Used Databricks to integrate easily with teh whole Microsoft stack.
  • Designed and mechanized Custom-constructed input connectors utilizing Spark, Sqoop and Oozie to ingest and break down informational data from RDBMS to Azure Data lake.
  • Involved in building an Enterprise Data Lake utilizing Data Factory and Blob storage, empowering different groups to work with more perplexing situations and ML solutions.
  • Involvement in working with Azure cloud stage (HDInsight, Databricks, Data Lake, Blob, Data Factory, Synapse, SQL DB and SQL DWH).
  • Broad involvement in working with SQL, with profound noledge on T-SQL (MS SQL Server).
  • Worked with data science group to do pre-processing and include feature engineering, helped Machine Learning algorithm in production.
  • Managed assets and scheduling over teh cluster utilizing Azure Kubernetes Service.
  • Performed information purging and applied changes utilizing Databricks and Spark information analysis.
  • Extensive information in Data changes, Mapping, Cleansing, Monitoring, Debugging, execution tuning and investigating Hadoop clusters.
  • Extensively utilized Databricks notebooks for interactive analysis utilizing Spark APIs.
  • Developed Spark Scala scripts for mining information and performed changes on huge datasets to handle ongoing insights and reports.
  • Developed a data pipeline using Kafka and Spark to store data into HDFS.
  • Developed spark applications in python (PySpark) on distributed environment to load huge number of CSV files with different schema in to Hive ORC tables.
  • Experience in Configure, Design, Implement and monitorKafkaCluster and connectors.
  • Used Azure Event Gridfor managing eventservice dat enables you to easily manage events across many differentAzureservices and applications.
  • Used Azure Synapse to oversee handling outstanding workloads and served data for BI and predictions.
  • Responsible for design & deployment ofSpark SQLscripts andScalashell commands based on functional specifications.
  • Facilitated information for interactive Power BI dashboards and reporting.
  • Data Ingestion to one or more Azure Services - (Azure Data Lake, Azure Storage, Azure SQL, Azure DW) and processing teh data in in Azure Databricks.
  • Using Linked Services/Datasets/Pipeline/ to extract, transform and load data from various sources such as Azure SQL, Blob storage, Azure SQL Data warehouse, write-back tool, and backwards, ADF pipelines were created.
  • Used Azure Synapse to bring these worlds together with a unified experience to ingest, explore, prepare, manage, and serve data for immediate BI and machine learning needs.
  • Worked onKafkaandSparkintegration for real time data processing.
  • Designed end to end scalable architecture to solve business problems using various Azure Components like HDInsight, Data Factory, Data Lake, Storage and Machine Learning Studio.
  • Developed JSON Scripts for deploying teh Pipeline in Azure Data Factory (ADF) dat process teh data using teh SQL Activity.
  • Scripting via Linux & OSX platforms: Bash, GitHub GitHub API.

Environment: Hadoop, Spark, Hive, Sqoop, HBase, Oozie, Talend, Kafka Azure (HDInsight, Databricks, Data Lake, Blob Storage, Data Factory, SQL DB, SQL DWH, AKS), Scala, Python, Cosmos DB, MS SQL, MongoDB, Ambari, PowerBI, Azure DevOps, Microservices, K-Means, KNN. Ranger, Git

Confidential, Nashville, TN

Data Engineer

Responsibilities:

  • Involved in writing optimizedPigScript along with developing and testingPig LatinScripts.
  • Involved in transforming data from Mainframe tables toHDFS, andHBasetables using Sqoop.
  • Acted for bringing in data underHBaseusing HBase shell alsoHBaseclient API.
  • Implemented teh workflows usingApache Oozieframework to automate tasks.
  • Designed and implemented Incremental Imports intoHivetables.
  • Visualized teh results using Tableau dashboards and teh Python Seaborn libraries were used for Data interpretation in deployment.
  • Involved in creatingHivetables, loading with data and writingHive queriesdat will run internally in MapReduce way
  • Involved in collecting, aggregating and moving data from servers to HDFS usingFlume.
  • Imported and Exported Data from Different Relational Data Sources like DB2, SQL Server, Teradata to HDFS usingSqoop.
  • Involved in data ingestion intoHDFSusingSqoopfor full load and Flume for incremental load on variety of sources like web server,RDBMSand Data API’s.
  • Collected data using Spark Streaming fromAWSS3bucket in near-real- time and performs necessary Transformations and Aggregations to build teh data model and persists teh data inHDFS.
  • InstalledOozieworkflow engine to run multipleHiveandPigjobs which run independently with time and data availability.
  • CreatedETLMapping with Talend Integration Suite to pull data from Source, apply transformations, and load data into target database.
  • Worked on POC for IOT devices data, with spark.
  • Stored teh time-series transformed data from teh Spark engine built on top of a Hive platform to Amazon S3 and Redshift.
  • Automatically scale-up teh EMR instances based on teh data.
  • Imported Bulk Data intoHBaseUsing MapReduce programs.
  • UsedSCALAto storestreaming datato HDFS and to implementSparkfor faster processing of data.
  • Worked on creating theRDD's,DF's for teh required input data and performed teh data transformations using Spark Python.
  • Involved in migrating tables fromRDBMSintoHivetables usingSQOOPand later generate visualizations using Tableau.

Environment: Hadoop, Cloudera, Flume, HBase, HDFS, MapReduce, AWS, YARN, Hive, Pig, Sqoop, Oozie, Tableau, Java, Solr.

Confidential

Hadoop Developer

Responsibilities:

  • Loaded data fromUNIXfile system to HDFS and writtenHive User Defined Functions.
  • Used Sqoop to load data from DB2 toHBasefor faster querying and performance optimization.
  • Worked on streaming to collect dis data fromFlumeand performed real time batch processing.
  • DevelopedHive scriptsfor implementing dynamic partitions.
  • Designed, developed and did maintenance of data integration programs in a Hadoop and RDBMS environment with both traditional and non-traditional source systems.
  • Developed suit of Unit Test Cases forMapper, Reducer and Driverclasses using testing library.
  • Collected teh logs data from web servers and integrated in to HDFS using Flume.
  • Worked on developingETL Workflowson teh data obtained using Scala for processing it in HDFS andHBaseusing Oozie.
  • Written ETL jobs to visualize teh data and generate reports from MySQL database using DataStage.
  • Worked extensively on AWS Components such as Airflow, Elastic Map Reduce (EMR), Atana, Snowflake.
  • DevelopedPythonscripts to find vulnerabilities with SQL Queries by doing SQL injection
  • DevelopedMapReducejobs in bothPIGandHivefor data cleaning and pre-processing.
  • Experience in writing SQOOP Scripts for importing and exporting data from RDBMS to HDFS.
  • Developed python code for different tasks, dependencies, SLA watcher and time sensor for each job for workflow management and automation using Airflow tool
  • Installed and configuredHive, Pig, Sqoop, Flume and Oozieon teh Hadoop cluster.
  • DevelopedSqoopscripts for loading data into HDFS from DB2 and pre-processed with PIG.
  • Automated teh tasks of loading teh data into HDFS and pre-processing with Pig by developing workflows using Oozie

Environment: Hadoop, HDFS, Hive, Pig, Flume, Mapper, Flume, ETL Workflows, HBase, Python, Sqoop, Oozie, DataStage, Linux, Relational Databases, SQL Server, DB2.

Confidential

Data Analyst

Responsibilities:

  • Gathered requirements from Business and documented for project development.
  • Prepared and maintained documentation for on-going projects.
  • Worked with Informatica Power Center for data processing and loading files.
  • Extensively worked with Informatica transformations.
  • Prepared ETL standards, Naming conventions and wrote ETL flow documentation for Stage, ODS and Mart.
  • Created data maps in Informatica to extract data from Sequential files.
  • Coordinated design reviews, ETL code reviews with teammates.
  • Interacted with key users and assisted them with various data issues, understood data needs and assisted them with Data analysis.
  • Worked with SQL*Loader tool to load teh bulk data into Database.
  • Developed mappings using Informatica to load data from sources such as Relational tables, Sequential files into teh target system.
  • Collect and link metadata from diverse sources, including relational databases and flat files.
  • Performed Unit, Integration and System testing of various jobs.
  • Extensively worked on UNIX Shell Scripting for file transfer and error logging.

Environment:Informatica Power Center, Oracle 10g, SQL Server, SQL*Loader, UNIX Shell Scripting, ESP job scheduler

We'd love your feedback!