Big Data Developer Resume
Austin, TX
SUMMARY
- Around 7 years of experience as a Data Engineer and extensively worked with designing, developing, and implementing Data models for enterprise - level applications and BI solutions.
- Experience in designing and building Data Management Lifecycle covering Data Ingestion, Data integration, Data consumption, Data delivery, and integration Reporting, Analytics, and System-System integration.
- Proficient in Big Data environment and Hands-on experience in utilizing Hadoop environment components for large-scale data processing including structured and semi-structured data.
- Strong experience with all phases including Requirement Analysis, Design, Coding, Testing, Support, and Documentation.
- Extensive experience with Azure cloud technologies like Azure Data Lake Storage, Azure Data Factory, Azure SQL, Azure Data Warehouse, Azure Synapse Analytical, Azure Analytical Services, Azure HDInsight, and Databricks.
- Experience with monitoring the web services using Hadoop and Spark for controlling the applications and analysing their operation and performance.
- Experienced in Python data manipulation for loading and extraction as well as with Python libraries such as NumPy, Pandas, and SciPy for data analysis and numerical computations.
- Good knowledge and experience with NoSQL databases like HBase, Cassandra, and MongoDB and SQL databases like Teradata, Oracle, PostgreSQL, and SQL Server.
- Experience in the development and design of various scalable systems using Hadoop technologies in various environments and analysing data using MapReduce, Hive, and PIG.
- Hands-on use of Spark and Scala to compare the performance of Spark with Hive and SQL, and Spark SQL to manipulate Data Frames in Scala.
- Strong knowledge in working with ETL methods for data extraction, transformation, and loading in corporate- wide ETL Solutions and Data Warehouse tools for reporting and data analysis.
- Hands-on experience in designing and implementing data engineering pipelines and analysing data using Hadoop ecosystem tools like HDFS, Spark, Sqoop, Hive, Flume, Kafka, Impala, PySpark, Oozie, and HBase.
- Strong experience with all phases including Requirement Analysis, Design, Coding, Testing, Support, and Documentation.
- Experience with different ETL tool environments like SSIS, Informatica, and reporting tool environments like SQL Server Reporting Services, and Business Objects.
- Experience in deployment of applications and scripting using the Unix/Linux Shell scripting.
- Solid knowledge of Data Marts, Operational Data Store, OLAP, Dimensional Data Modelling with Star Schema Modelling, Snowflake Modelling for Dimensions Tables using Analysis Services.
- Extensive experience with various databases like Teradata, MongoDB, Cassandra DB, MySQL, Oracle, and SQL Server.
- Participated in the design and development of data solutions that use Azure Data Factory and Azure Synapse Analytics to support the Business Intelligence team to deliver actionable insights.
- Worked with the Data Science team to construct Gold tables based on business requirements for premium members providing further athlete performance and analysis and created managed Hive tables in Databricks.
- Experience in Creating Teradata SQL scripts using OLAP functions like rank and rank over to improve the query performance while pulling the data from large tables.
- Strong Experience in working with Databases like Teradata and proficiency in writing complex SQL, PL/SQL for creating tables, views, indexes, stored procedures, and functions.
- Knowledge and experience with Continuous Integration and Continuous Deployment using containerization technologies like Docker and Jenkins.
- Excellent working experience in Agile/Scrum development and Waterfall project execution methodologies
TECHNICAL SKILLS
Hadoop/Big Data ecosystems: HDFS, Map Reduce, Sqoop, Flume, Pig, Hive, Oozie, Impala, Zookeeper and Cloudera Manager, Zookeeper, Spark, Scala
NoSQL Database: HBase, Cassandra
Tools: and IDE Eclipse, NetBeans, Toad, Putty, Maven, DB Visualizer, VS Code, Qlik Sense, Qlik View
Languages: SQL, PL/SQL, JAVA, Scala, Python
Databases: Oracle, SQL Server, MySQL, DB2, PostgreSQL, Teradata
Tracking Tools and Control: SVN, GIT, Maven
ETL Tools: OFSAA, IBM DataStage
Cloud Technologies: Azure, AWS
PROFESSIONAL EXPERIENCE:
Confidential, Austin, Tx
Big Data Developer
Responsibilities:
- Research and recommend a suitable technology stack for Hadoop migration considering current enterprise architecture.
- Expertise in writingHadoopJobsfor analyzing structured and unstructured data usingHDFS,Hive, HBase, Pig, Spark, Kafka, Scala, Oozie, andTalend ETL.
- Extensively usedSparkstack to develop preprocessing job which includes RDD, Datasets, and Data frame APIs to transform the data for upstream consumption.
- Build and maintain scalable data pipelines using theHadoop ecosystemand other open-source components likeHive, andHBase.
- Good understanding ofNoSQL Databases and hands-on work experience in writing applications on No SQL databases likeCassandraandMongo DB.
- Good knowledge in querying data fromCassandrafor searching grouping and sorting.
- Involved in various NoSQL databases likeHBase, Cassandrain implementing and integration.
- Involved in developing of JDBCDAOs and DTOs, access of advanced SQL and PL/SQL stored procedures on database systems using spring templates.
- Onboarding of new data into Splunk. Troubleshooting Splunk and optimizing performance.
- Developed multiple MapReduce jobs in java for data cleaning and pre-processing.
- Worked on extracting and enriching relational databases data between multiple tables using joins inSpark.
- Worked on writing APIs to load the processed data to Hive tables.
- ExperiencedScalain using andspark streamingandAkkafor ongoing transactions for customers.
- Replaced the existing Map Reduce programs intoSparkapplication using Scala.
- Developed the Hive UDF's to handle data quality and create filtered datasets for further processing
- Experienced in writing Sqoop scripts to import data into Hive/HDFS from RDBMS.
- Applied Azure Data Factory to ingest raw data from several data sources and store it in Azure Data Lake Storage Gen2 for downstream processing.
- Created dynamic pipeline with stored procedures or data flow in Azure Data Factory to update different dimensional tables for reducing the number of activities and easier maintenance.
- Used Azure Databricks (PySpark) to exact raw data and transform it into bronze by adding metadata, then cleansed, validated, and enriched bronze data to silver/gold tables and persisted them in ADLS gen 2 based Delta Lake
- Ingested real-time streaming sports game pipelines with Event Hubs and processed the athlete score and performance in real-time by Azure Stream Analytics for providing simple data to software engineers, which is displayed on the mobile application.
- Conducted daily batch incremental load by Azure Data Factory to update the data for customer use. · Implemented Slowly Changing Dimension Type 2 in Delta Lake as the update strategy to keep the full historical player and team profiles.
- Worked with the Data Science team to construct Gold tables based on business requirements for premium members providing further athlete performance and analysis and created managed Hive tables in Databricks
- SupportedMap Reduce Programsthat are running on the cluster and wroteMapReduce jobsusingJava API.
- Optimized Hive QL scripts by using execution engines like Tez, Spark.
- Developed Hive queries to analyze the data in HDFS to identify issues and behavioral patterns.
- Able to use Python Pandas, NumPy modules for Data analysis, Data scraping, and parsing.
- Deployed applications using Jenkins’s framework integrating Git- version control with it.
- Participated in production support regularly to support the Analytics platform
- Participated in the design and development of data solutions that use Azure Data Factory and Azure Synapse Analytics to support the Business Intelligence team to deliver actionable insights.
- Used Git and Azure DevOps to perform version control, peer review, and enhance collaboration.
- Worked under Agile Scrum methodology with 2-week sprints, and participated in stand-up meetings, sprint planning, retrospective meetings, and code reviews
Environment: Hadoop, HDFS, AWS, Hive, BigQuery, Spark SQL, GCP, MapReduce,Sparkstreaming, Abinitio, Sqoop, Oozie, Jupiter Notebook, Docker, Kafka,Spark, Scala, Talend, Shell Scripting.
Confidential
Big Data Developer
Responsibilities:
- Performed advanced procedures like text analytics and processing using the in-memory computing capabilities ofSpark.
- DevelopedSparkcode using Scala andSpark-SQL for faster processing and testing.
- Worked onSparkSQL for joining multi hive tables and write them to a final hive table.
- Configured Spark Streaming to receive real-time data from the Apache Kafka and store the stream data toHDFSusingScala.
- Worked on scalable distributed data system using Hadoop ecosystem in AWS EMR and MapR (MapR data platform).
- Developed prototypeSparkapplications usingSpark-Core, Spark SQL, Data Frame APIand developed several custom User-defined functions inHive & Pig using Java & python
- Importing the data intoSparkfromKafkaConsumer group usingSpark Streaming APIs.
- Good Knowledge of reporting and data visualization tools like Oracle Data Visualization Desktop, Tableau, and Grafana.
- Worked on Google Cloud Platform (GCP) Services like Vision API, Instances.
- Used Informatica Cloud Data Integration for global, Data Cloud Architect, Azure, Ansible, Jenkins, Docker, Kubernetes, DevOps, Automation, CI/CD, distributed data warehouse, and analytics projects.
- Utilized Azure Data Factory to create, Data Cloud Architect, Azure, Ansible, Jenkins, Docker, Kubernetes, DevOps, Automation, CI/CD, schedule and manage data pipelines.
- Developed a POC for project migration from on prem Hadoop MapR system to GCP/Snowflake.
- Build cluster on AWS environment using EMR using S3, EC2, and Redshift.
- Built dashboards and visualizations on top of MapR-DB and Hive using Oracle data visualizer desktop. Built real-time visualizations on top of Open TSDB using Grafana.
- Used DataStax Cassandra along with Pentaho for reporting.
- Designed, configured, and deployed Amazon Web Services (AWS) for a multitude of applications utilizing the AWS stack (Including EC2, Glue, Data pipeline EMR, SNS, S3, RDS, Cloud Watch, SQS, IAM), focusing on high-availability, fault tolerance, and auto-scaling.
- Worked with AWS Glue jobs to transform data to a format that optimizes query performance for Athena.
- Working on new system architecture to replace client's current crediteTradingplatform usingFlink.
- ImplementedSparkRDD transformations to Map business analysis and apply actions on top of transformations.
- CreatedSparkjobs to do lighting speed analytics over theSparkcluster.
- EvaluatedSpark's performance vs Impala on transactional data.
- Experienced in developing scripts for doing transformations usingScala.
- Experienced in creating data pipeline integratingKafkawithspark streamingapplication usedScalafor writing applications.
- Developed customizedUDFs in Javafor extending Pig and Hive functionality.
- Usedspark SQLfor reading data from external sources and processes the data usingScalacomputation framework.
- UsedSparktransformations and aggregations to perform min, max, and average on transactional data.
- Extracted files from databases through Sqoop and placed in HDFS and processed through spark.
- Experienced in migrating Hive QL into Impala to minimize query response time.
- Experience using Impala for data processing on top of HIVE for better utilization.
- Developed and optimizedPigandHiveUDFsto implement the methods and functionality of Javaas required.
- Wrote queries UsingCassandra CQLto create, alter, insert and delete elements.
- Hands-on experience working onNoSQLdatabases includingHBase,MongoDB,Cassandra, and its integration withHadoop cluster.
- ImplementedAWS EC2, Key Pairs, Security Groups, Auto Scaling, ELB, SQS, and SNSusingAWS APIand exposed as theRestful Web services andimplementedReporting, Notification servicesusingAWS API.
- Performed querying of both managed and external tables created by Hive using Impala.
- Developed Impala scripts for end-user/analyst requirements for Adhoc analysis.
- Continuous monitoring and managing the Hadoop cluster through Cloudera Manager.
- UsedSparkAPI over Cloudera Hadoop YARN to perform analytics on data in Hive.
- UsedHTML, CSS, XML, JavaScriptandJSPfor interactive cross browser functionality and complex user interface.
- Experience in writingcustom UDFsfor Hive to in corporate methods and functionality of Java into andHQLHIVESQL.
- Collected data usingSparkStreaming from AWSS3 bucket in near-real-time and performs necessary Transformations and Aggregations to build the data model and persists the data in HDFS.
- Responsible for creating Hive tables, loading with data, and writing Hive queries.
- Optimized Hive QL by using execution engines like Tez,Spark.
- Responsible for creating mappings and workflows to extract and load data from relational databases, flat file sources, and legacy systems using Abinitio.
- Fetch and generate monthly reports, Visualization those reports using Tableau.
- Used Oozie Workflow engine to run multiple Hive jobs.
Environment: Hadoop, Cloudera, Flume, HBase, GCP, HDFS, MapReduce, YARN, Hive, Sqoop, Oozie, Tableau, Abinitio, JUnit, agile methodologies, UNIX
Confidential
Hadoop/ETL Developer
Responsibilities:
- Implemented ETL Abinitio designs and processes for a load of data from the sources to the target warehouse.
- Experience in developing scalable real-time applications for ingesting clickstream data using Kafka Streams and Spark Streaming.
- Configured several nodes EC2 Hadoop cluster to transfer the data from S3 to HDFS and vice-versa and to direct input and output to the Hadoop MapReduce framework.
- Experience with different table structure, file formats, partitioning and bucketing concepts in Hive.
- Experienced in converting Hive scripts into spark using Scala and optimized spark jobs.
- Developed ETL Processes in AWS Glue to migrate Campaign data from external sources like S3 ORC/Parquet/Text Files into AWS Redshift.
- Experience in developing Spark applications using Spark-SQL in Databricks for data extraction, transformation, and aggregation from multiple file formats for Analyzing & transforming the data to uncover insights into the customer usage patterns.
- Create PySpark frame to bring data from DB2 to Amazon S3.
- Pushed application logs and data streams logs to Kibana server for monitoring and alerting purpose.
- Developed optimized and tuned ETL operations in Hive and Spark scripts using techniques such as partitioning, bucketing, vectorization, serialization, configuring memory and number of executors.
- Implemented cloud integrations to GCP and Azure for bi-directional flow setups for data migrations.
- Experience in developing Spark applications using Spark-SQL in Databricks for data extraction, transformation, and aggregation from multiple file formats for Analyzing & transforming the data to uncover insights into the customer usage patterns.
- Pushed application logs and data streams logs to Kibana server for monitoring and alerting purpose.
- Experience designing solutions in Azure tools like Azure Data Factory, Azure Data Lake, SQL DWH, Azure SQL & Azure SQL Data Warehouse, Azure Functions.
- Worked on migrating data from HDFS to Azure HD Insights and Azure Databricks.
- Migrated existing processes and data from our on-premises SQL Server and other environments to Azure Data Lake.
- Developed and Tuned Spark Streaming application using Scala for processing data from Kafka.
- Developed Jenkins pipelines for continuous integration and deployment purpose.
- Experience in working on analyzing snowflake datasets performance.
- Worked on building pipelines using snowflake for extensive data aggregations.
Environment: Kafka, Spark, Sqoop, Hive, Azure, Databricks, Grafana, Jenkins, Azure Data Lake, Azure SQL, Jenkins, Grafana, Python, Shell, Microservices, Restful API's
