We provide IT Staff Augmentation Services!

Hadoop/bigdata Engineer Resume

3.00/5 (Submit Your Rating)

Dallas, TX

PROFESSIONAL SUMMARY:

  • Around 7 years of professional IT experience with Java , Python, Elastic search, Big data Environment , Hadoop Ecosystem and good experience in Spark, SQL, Java Development.
  • Worked as a Architect i.e by being bridge between Data Scientists, Engineers, and the organizational needs.Worked with a team of 5 in almost all the projects.
  • Worked as Technical Lead with Elastic search , Java , Hadoop, Spark, Scala, Hive.
  • Worked as a Technical Architect with Data Analytic , Hadoop, Java and Scala.
  • Hands on experience across Hadoop Eco System that includes extensive experience in Big Data technologies like HDFS, MapReduce, YARN, Spark, Apache spark Sqoop, Hive, Pig, Impala, Oozie, Oozie Coordinator, schema registry, KSQL, Rest proxy, Replicator, ADB, Operator and Kafka Control center, Zoo - Keeper and Apache Cassandra, HBase.
  • Experience in using various tools like Sqoop, Flume, Kafka, NiFi, Pig to ingest structured, semi-structured and unstructured data into the cluster.
  • Hands on experience in JAVA 8 and DEVOPS.
  • Working knowledge in AWS environment and AWS connect, Cassandra,spark with Strong experience in Cloud computing platforms such as AWS services.
  • Designing both time driven and data driven automated workflows using Oozie and used Zookeeper for cluster co-ordination.
  • Expertise in TALEND
  • Experience in Hadoop cluster using Cloudera's CDH, Horton works HDP.
  • Hands on experience in SPRING BOOT and SPRING SECURITY .
  • Developed highly optimized Spark applications to perform various data cleansing, validation, transformation and summarization activities according to the requirement
  • Data pipeline consists Spark, Hive and Sqoop and custom build Input Adapters to ingest, transform and analyze operational data.
  • Real time experience in HL7.
  • Hands on experience on Aurora Postgre.
  • Hands on experience with R, API, ETL, DATA STAGE, BIG QUERY .
  • Developed Spark jobs and Hive Jobs to summarize and transform data.
  • Involved in converting Hive/SQL queries into Spark transformations using Spark Data Frames and Python.
  • Hands on experience on KUBERNETES.
  • Hands on experirnce in Telecom Domain.
  • Hands on experience with COBOL
  • Experience in working with structured data using HiveQL, join operations, Hive UDFs, partitions, bucketing and internal/external tables.
  • Expertise in SOLARIS
  • Experienced in Ranger, Sentry, RBAC, role types, Control plane, Data plane.
  • Expertise in writing Map-Reduce Jobs in Java, Python for processing large sets of structured, semi-structured and unstructured data sets and stores them in HDFS.
  • Experience working with Python, UNIX and shell scripting.
  • Experience in CONDA environment.
  • Hands on experience with Cassandra.
  • Experience in Extraction, Transformation and Loading ( ETL ) of data from multiple sources like Flat files and Databases.
  • Good knowledge of cloud integration with AWS using Elastic Map Reduce ( EMR ), Simple Storage Service (S3), EC2, Redshift and Microsoft Azure.
  • Hands on Experience with MongoDB and stream sets.
  • Experienced in DATA ANALYSIS, DATA MINING.
  • Experience with complete Software Development Life Cycle ( SDLC ) process which includes Requirement Gathering, Analysis, Designing, Developing, Testing, Implementing and Documenting.
  • Worked with waterfall and Agile methodologies.
  • Good team player with excellent communication skills with strong attitude towards learning new technologies.
  • Hands on Experience in Spark architecture and its integrations like Spark SQL , Data Frames and Datasets APIs .
  • Worked on Spark for enhancing the executions of current processing in Hadoop utilizing Spark Context , Spark SQL, Data Frames and RDD’s .
  • Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs, Spark SQL and Python.
  • Hands on experience Using Hive Tables by Spark , performing transformations and Creating Data Frames on Hive tables using Spark.
  • Used Spark-Structured-Streaming to perform necessary transformations.
  • Expertise in converting Map Reduce programs into Spark transformations using Spark RDD's

TECHNICAL SKILLS:

HADOOP: HDFS, MapReduce, Hive, beeline, Sqoop, Flume, Oozie, Impala, pig, Kafka, Zookeeper, NiFi, Cloudera Manager, Horton Works

Spark Components: Spark Core, Spark SQL (Data Frames and Dataset), Scala, Python, Apache spark

Cloud Technologies: Amazon Web Services(AWS)

Programming Languages: Core Java, Scala, Shell, Hive-QL, Python

Web Technologies: HTML, JQuery, Ajax, CSS, JSON, JavaScript.

Operating Systems: Linux, Ubuntu, Windows 10/8/7

Databases: Oracle, MySQL, SQL Server,

NoSQL Databases: Hbase, Cassandra, MongoDB

Cloud: AWS Cloud Formation, Azure

Version controls and Tools: GIT, Maven, SBT, CBT

Methodologies: Agile, Waterfall

IDES & Command Line Tools: Eclipse, Net Beans, IntelliJ

WORK EXPERIENCE:

Confidential, DALLAS, TX

Hadoop/BigData Engineer

Responsibilities:

  • Worked with product owners, Designers, QA and other engineers in Agile development environment to deliver timely solutions to as per customer requirements.
  • Transferring data from different data sources into HDFS systems using Kafka producers, consumers and Kafka brokers.
  • Experienced in writing live Real-time Processing using Spark Streaming with Kafka on AWS EMR, EDL, AWS Connect.
  • Data Analytics helps to provide technical support to the client and train the end users for the final product
  • Used Spark API over EMR Cluster Hadoop YARN to perform analytics on data .
  • Played key role in Migrating Teradata objects into SnowFlake environment.
  • LUNGI is in-memory databases provide a predictive response time which is suited to real time and near real time applications
  • Bench marked Cassandra cluster based on the expected traffic for the use case and optimized for low latency
  • Troubleshoot read/write latency and timeout issues in CASSANDRA Installation, Talend, Configuration, Upgrade, patching of Oracle.
  • Experience with Snowflake Multi-Cluster Warehouses .
  • Managed Docker orchestration and Docker containerization using Kubernetes.
  • Used Kubernetes to orchestrate the deployment, scaling and management of Docker Containers.
  • Aurora with Postgre helps in Configuring the parameter group by managing IP traffic using a security group, Auditing the database log files also maintenance and management activities
  • Aurora with Postgre helps in Planning backup and recovery strategies, User management and Monitoring the database
  • Participate in the design and development of features related to the integration of the Istio framework within Red Hat OpenShift Develop.
  • Ambassdor helps in Implementing Spring boot microservices to process the messages into the Kafka cluster setup. Worked as Onshore lead to gather business requirements and guided the offshore team on
  • Data Analytics demonstrate the features of Istio and its use within Red Hat OpenShift Provide documentation in support of the Istio integration within Red Hat OpenShift
  • Administer and provide day to day engineering and operational support for the Azure environment.
  • Azure provide second and third level technical support for a mission critical corporate Windows server environment including.
  • Google Cloud Platform team helps customers transform and evolve their business through the use of Google's global network, Azure Data, web-scale data centers and software infrastructure.
  • As part of the Public Sector Engineering team, you will shape the future of public sector agencies by helping customers deploy cloud-based solutions.
  • For AWS Developing excellent quality software using agile techniques such as Test Driven Development and Pair Programming
  • Scala actively contribute in development, Confidential planning, support ( second line) & release management.
  • Considering the Dev part of the DevOps, Talend, software developers and QA engineers are at the very heart of the organization. Without them, you simply don't really need DevOps. when it comes to the HL7 protocol, it is used in a host of projects. All the apps or programs that need medical device interoperability, they prefer to use HL7.
  • Mappings can be handled by Administrator within Airflow UI when the Authentication backend doesn’t provide the Group info for the authenticated user.
  • This is done to account for any modifications made to user-group mapping within the authentication backed since the user’s last login to Airflow.
  • By using scala assisting the PM and Technical Lead in the project planning process, provide detailed work estimates.
  • Stream Sets helps in Collaborating with internal SDR, Marketing, Solution Engineering, Products and broader Field organization teams to identify and strategically target new, high-value opportunities within their respective territories
  • Stream sets happens to Cultivating sales by nurturing inbound leads and leveraging a multi-channel, multi-touch approach, Talend applying an outbound prospecting strategy, identifying personas, compelling events and business drivers in various verticals.
  • Develop new and existing modules in Scala while working with developers across the globe.
  • MongoDB grants access to data and commands through role-based authorization and provides built-in roles that provide the different levels of access commonly needed in a database system.
  • By using MongoDB You can additionally create user-defined roles. A role grants privileges to perform sets of actions on defined resources.
  • Scala helps in assisting the PM, the Technical Lead and Technical Architect in the project planning process, provide detailed work estimates.
  • Experience with Snowflake Virtual Warehouses.
  • Experienced in Data Bricks
  • Testing Telecom Domain with Sample OSS/BSS Test cases.
  • Developed highly optimized Spark applications to perform various data cleansing, validation, transformation and summarization activities according to the requirement
  • Azure Data consists Spark, Hive and Sqoop and custom build Input Adapters to ingest, transform and analyze operational data.
  • Here we use Spark SQL that helps in broadcast join (aka broadcast hash join) instead of hash join to optimize join queries
  • Played key role in Migrating Teradata objects into SnowFlake environment.
  • Experience with Snowflake Multi-Cluster Warehouses .
  • Experience with Snowflake Virtual Warehouses.
  • Broadcast join can be very efficient for joins between a large table (fact) with relatively small tables (dimensions) that could then be used to perform a star-schema join.
  • CXSQL was configured to work as transparent cache for database management systems such as MySQL and PostgreSQL
  • Data Stage extracts, transforms, and loads data from source to the target.
  • Data Bricks is an integrated set of tools for designing, developing, running, compiling, and managing applications
  • CONDA is used to customize the environment and for instealling the extensions with Angular JS.
  • It is a definition set out in Graph Theory, but DAGs have been used, without a formal definition for thousands of years before Graph Theory was formalised.
  • Here DAG is used for the formation of represantation with graph.
  • Developed Spark jobs and Hive Jobs to summarize and transform data.
  • 8 is used here for understanding the threading/concurrency by including synchronized blocks, wait notify also executors, threadpools, fork/join, blocking queues, semaphores, countdown etc.
  • In AZILE the internal code quality is the main focus in this quadrant, and it consists of test cases which are technology driven and are implemented to support the team, it include
  • 1. Unit Tests
  • 2.Component Tests
  • Involved in converting Hive/SQL queries into Spark transformations using Spark Data Frames and Python.
  • Used Oozie for automating the end-to-end data pipelines and Oozie coordinators for scheduling the workflows.
  • Involved in creating Hive tables , loading data and writing hive queries, views and worked on them using Hive QL .
  • CLOUDERA is involved here in few initiatives, including helping establish a cloud-based collection and analysis tool that identified suspected human tracficking networks and indivisuals.
  • A set is a collection where you store distinct values while a seq is a more generalized version of collection its common superclass for some often used collections like Array,List etc
  • Performed Optimizations of Hive Queries using Map side joins, dynamic partitions and Bucketing.
  • Applied Hive queries to perform data analysis on HBase using the serde tables in meeting the data requirements for the downstream applications.
  • Responsible for executing hive queries using Hive Command Line, Ranger, Sentry, Lungi, RBAC, role types, Control plane, EDL, Data plane, Web GUI HUE and Impala to read, write and query the data into HBase.
  • Talend is participated in all phases of development life-cycle with extensive involvement in the definition and design meetings, functional and technical walkthroughs.
  • Data Bricks is an integrated set of tools for designing, developing, running, compiling, and managing applications
  • Created Talend jobs to copy the files from one server to another and utilized Talend FTP components.
  • Map type here are used for indicating the spread out on something that are shown for distributed areas.
  • Elastic App Search provides a rich set of APIs for ingestion and searching content along with an intuitive UI for analyzing and tuning relevance, and an open-source library to quickly implement rich search experiences when it comes to the HL7 protocol, it is used in a host of projects. All the apps or programs that need medical device interoperability, they prefer to use HL7.
  • Implemented Map R educe secondary sorting to get better performance for sorting results in MapReduce programs.
  • Troubleshoot Issues. When technical issues with the product arise, production support engineers must act quickly to analyze the available data and find the root cause of the problem
  • Messaging storing via web applications is done with the help of SPRING BOOT.
  • Authentication and Authorization is held by sourcing SPTING BOOT.
  • Worked ETL phases of the data like data cleansing, Angular JS, data massaging and data cleanup, filtering the data which is useful for model building.
  • Designed ETL workflows on Tableau, Deployed data from various sources to HDFS.
  • Worked on data engineer, data warehousing and ETL tools like Informatica, Talend, and Pentaho.
  • Load and transform large sets of structured, semi structured that includes Avro, sequence files.
  • Production support helps to participate in all stages of the product development process, including designing, building, and testing. They also create useful tools such as internal software to automate key processes or platforms where customers can send inquiries and review.
  • Lead effort to develop technical standards to support and operate technologies within the current or future Microsoft azure.
  • AKKA is used for language bindings that exist for both JAVA and SCALA, and also used for multiple programming models for concurrency.
  • Microsoft Azure IaaS Monitoring and Management, manage and monitor IaaS deployments by Log Analytics and Log Search to “drill down” into the most important data in your IaaS systems.
  • Worked on migration of all existed jobs to Spark, to get performance and decrease time of execution.
  • Implemented usage of Amazon EMR for processing Big Data across a Hadoop Cluster of virtual servers on Amazon Elastic Compute Cloud ( EC2 ) and Amazon Simple Storage Service ( S3 ).
  • Using Hive join queries to join multiple tables of a source system and load them to Elastic search tables.
  • Experience with ELK Stack in building quick search and visualization capability for data.
  • Experience with different data formats like Json, Avro, parquet, ORC formats and compressions like snappy & bzip.
  • Coordinated with the testing team for bug fixes and created documentation for recorded data, agent usage and release cycle notes.

Environment : Hadoop, Azure Data, Big Data, Spring boot, Production Support, Data engineer, Talend, HDFS, Data stage, EDL, Stream sets, Data Bricks, Elastic Search, MongoDB, AWS Connect, Cassandra, Ranger, HL 7, Sentry, RBAC, role types, Control plane, Airflow, DAG, Data plane, Open Source, Solaris, Cobol, Telecom Domain, Snowflake, Google Cloud Platform, Scala, ETL, Java 8, AKA Frame work, Kubernetes, Lungi, .NETl, Azile, Ambassdor, Istio, Cassandra, API, SQL, Python, Oozie, Devops, Elastic search, NIFI, Hive, Broadcast joint, HBase, Aurora with Postgre, Big query, ETL, NiFi, Impala, Spark, Conda, Angular JS, Cloudera, Data analysis, AWS, Linux, Apache spark.

Confidential, Nashville, TN

Hadoop/Java/ETL Developer

Responsibilities:

  • Developed an EDW solution, which is a cloud based EDW and Data Lake that supports Data asset management, Data Integration, and continuous data analytic discovery workloads.
  • Developed and implemented real-time data pipelines with Spark Streaming, Kafka , and Cassandra to replace existing lambda architecture without losing the fault-tolerant capabilities of the existing architecture.
  • The data will be transferred from one integration form to another by using CLOUDERA platform.
  • Scala actively contribute in development, Confidential planning, support ( second line) & release management and work with support to investigate problems, perform root-cause analysis and deliver bug-fixes
  • Aurora with Postgre helps in Configuring the parameter group by managing IP traffic using a security group, Auditing the database log files also maintenance and management activities
  • Experienced in writing live Real-time Processing using Spark Streaming with Kafka on AWS EMR.
  • For AWS Developing excellent quality software using agile techniques such as Test Driven Development and Pair Programming
  • Scala gets transparent performance management process and Test own code and peer-test other developers’ code.
  • MongoDB provides a number of built-in roles that administrators can use to control access to a MongoDB system.
  • Bench marked Cassandra cluster based on the expected traffic for the use case and optimized for low latency
  • Troubleshoot read/write latency and timeout issues in CASSANDRA Installation, Configuration, Upgrade, patching of Oracle.
  • MongoDB allows us to create a new user within the system in a very efficient way. If the user which we are going to insert already exists, then it will return an error in response. If there is no user that already exists then insert that record within the system
  • According to the article Map-Side Join in Spark, broadcast join is also called a replicated join (in the distributed system community) or a map-side join (in the Hadoop community).
  • DataStage is an integrated set of tools for designing, developing, running, compiling, and managing applications
  • Created a Spark Streaming application to consume real-time data from Kafka sources and applied real-time data analysis models that we can update on new data in the stream as it arrives.
  • Production support entails interacting with product users, often external customers but sometimes also employees. These interactions can occur in various setups, including in-person meetings, phone calls, emails, and live messaging chats. In all of these cases, it’s vital to address concerns promptly and maintain a helpful attitude.
  • Successfully securing new business while identifying and articulating upsell opportunities and Partnering with StreamSets’s Customer Success team to analyze customer health, identify gaps and opportunities for upsell and cross sell, maximizing renewals.
  • Using Elastic search tables helps hive join queries to join multiple tables of a source system and load them.
  • Because production support engineers deal with product issues firsthand, they can readily suggest overall product improvements, such as features that customers want. Ideally, they should also proactively evaluate engineering processes and provide recommendations to increase efficiency.
  • .NET helps to ensure the compatibility of websites with newer software and operating system versions
  • .NET helps to understand the software lifecycle and determine
  • In-depth knowledge of. Snowflake Database, Schema and Table structures.
  • Experience in using Snowflake Clone and Time Travel.
  • In-depth understanding of NiFi.
  • Effective execution of DevOps strategy requires a combined effort of different individuals. CTOs are advised to pick a strong team of flexible and cooperative individuals with an average size of DevOps team being five people.
  • Experience in building ETL pipelines using NiFi.
  • Used Talend most used components (tMap, tDie, tConvert Type, tFlow Meter, tLog Catcher, tRow Generator, tSet Global Var, tHash Input & tHash Output and many more
  • The Google Cloud Platform team helps customers transform and evolve their business through the use of Google's global network, web-scale data centers and software infrastructure.
  • Worked on importing, transforming large sets of structured semi-structured and unstructured data.
  • Used Spark-Structured-Streaming to perform necessary transformations and data model which gets the data from Kafka in real time and Persists into HDFS.
  • Implemented the workflows using the Apache Oozie framework to automate tasks. Used Zookeeper to co-ordinate cluster services.
  • This is done to account for any modifications made to user-group mapping within the authentication backed since the user’s last login to Airflow.
  • SPRING BOOT is used here for frame work secruring java based applications at various layers with flexibility.
  • SPRING BOOT also provides Spring security for database authentication in LDAP, Ranger, Sentry, RBAC, role types, Control plane, Data plane, JAAS etc.
  • JAVA 8 helps in unit testing like JUNIT or TESTNG and .mocking framework like bmoc or easy mock.
  • Used Kubernetes to orchestrate the deployment, scaling and management of Docker Containers
  • Created various hive external tables, staging tables and joined the tables as per the requirement.
  • Implemented static Partitioning , Dynamic partitioning and Bucketing in Hive using internal and external table. Created Map side Join, Parallel Execution for optimizing the Hive queries.
  • AKKA emphasizes actor - based concurrency, with inspiration drawn from Erlang.
  • Most of the analysis tools used in here are from the Cloud based data.
  • For maintaining security and trafficking networks also CLOUD platform is used .
  • Developed and implemented hive, AZURE and spark custom UDFs involving date Transformations such as date formatting and age calculations as per business requirements.
  • DAGs are more modern, with multiple uses related to computer science. Data processing networks, version history and some data compression algorithms all utilize DAGs.
  • Written Programs in Spark using Scala and Python for Data quality check.
  • Deploying Azure VMs (Windows Server and Linux) in a highly available environment.
  • Written transformations and actions on Data Frames Data Bricks , used Spark SQL on data frames and data bricks to access hive tables into spark for faster processing of data.
  • Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs , Python and Scala.
  • Several methods throw a No Such Element Exception when no items exist in the invoking map.
  • A ClassCastException is thrown when an object is incompatible with the elements in a map.
  • A Null Pointer Exception is thrown if an attempt is made to use a null object and null is not allowed in the map.
  • Used Spark optimizations techniques like Cach e /Refresh tables, b roadcasting variables, Coalesce /Repartitioning, increasing memory overhead limits, handling parallelism and modifying the spark default configuration variables for performance tuning.
  • Performed various benchmarking steps to optimize the performance of Spark jobs and thus improve the overall processing.
  • Worked in Agile environment in delivering the agreed user stories within the Confidential time .

Environment : Hadoop, Big query, EMR, HDFS, Data analysis, Data engineer, Azure, Data Mining, CSQL, Data Bricks, Google Cloud Platform, Azile, Kubernetes, Datastage, AKA frame work, Broadcast joint, Hive, Java8, Talend, MongoDB, Devops, DAG, Sqoop, Elastic search, R, .NET, Snowflake, Cassandra, Angular JS, Ranger, Sentry, RBAC, role types, Stream sets, Airflow, Control plane, Data plane, Production support, Cassandra, Map types, Telecom Domain, API, ETL, Aurora with Postgre, Party kit, Solaris, Spring boot, Oozie, Condo, Spark, Python Scala, Kafka, Python, Cloudera, Linux.

Confidential, Bowie, Maryland

Hadoop Developer

Responsibilities:

  • Responsible for building scalable distributed data solutions using Hadoop cluster environment with HortonWorks distribution.
  • Used Sqoop to load the data from relational databases.
  • Involved in converting Hive/SQL queries into spark transformations using Spark RDD’s.
  • Worked with CSV, Jason, Avro and Parquet file formats.
  • Elastic App Search provides a rich set of APIs for ingestion and searching content along with an intuitive UI for analyzing and tuning relevance, and an open-source library to quickly implement rich search experiences
  • Implemented usage of Amazon EMR for processing Big Data across Hadoop Cluster of virtual servers on Amazon Elastic Compute Cloud (EC2) and Amazon Simple Storage Service(S3).
  • Worked on Kafka to collect and load the data on Hadoop file systems.
  • Used Hive to form an abstraction on top of structured data resides in HDFS and implemented Partitions , Buckets on HIVE tables.
  • Developed and implemented real-time data pipelines with Spark Streaming.
  • Designed, developed data integration programs in a Hadoop environment with NoSQL data store HBase for data access and analysis.
  • Worked with Python , to develop analytical jobs using PySpark API of spark.
  • Using Job management scheduler apache Oozie to execute the workflow.
  • Using Ambari to monitor node’s health, status of the jobs and to run the analytics jobs in Hadoop clusters.
  • Experience with pyspark for using spark libraries by using python scripting for data analysis.
  • Worked on Tableau to build customized interactive reports, worksheets, and dashboards.
  • Involved in performance tuning of spark jobs using Cache and by utilizing complete advantage of cluster environment.

Environment: Hadoop, Spark, Scala, Python, Talend, Kafka, Hive, Sqoop, Pyspark, Ambari, Oozie, HBase, Tableau, Jenkins, Horton Works.

Confidential

Java Developer

Responsibilities:

  • Designed and developed Web Services using Java/J2EE in WebLogic environment. Developed web pages using Java Servlet, JSP, CSS, Java Script, DHTML, and HTML . Added extensive Struts validation. Wrote Ant scripts to build and deploy the application.
  • Involve in the Analysis, Design, and Development and Unit testing of business requirements.
  • Developed business logic in JAVA/J2EE technology .
  • Implemented business logic and generated WSDL for those web services using SOAP .
  • Worked on Developing JSP pages
  • Implemented Struts Framework .
  • Developed Business Logic using Java/J2EE.
  • Modified Stored Procedures in Oracle Database.
  • Developed the application using Spring Web MVC framework.
  • Worked with Spring Configuration files to add new content to the website.
  • Worked on the Spring DAO module and ORM using Hibernate. Used Hibernate Template and Hibernate Dao Support for Spring-Hibernate Communication.
  • Configured Association Mappings such as one-one and one-many in Hibernate
  • Worked with JavaScript calls as the Search is triggered through JS calls when a Search key is entered in the Search window
  • Worked on analyzing other Search engines to make use of best practices.
  • Collaborated with the Business team to fix defects.
  • Worked on XML , XSL and XHTML files.
  • As part of the team to develop and maintain an advanced search engine , would be able to attain

Environment : Java 1.6, J2EE, Talend, Eclipse SDK 3.3.2, Java Spring 3.x, jQuery, Oracle 10i, Hibernate, JPA, Json, Apache Ivy, SQL, stored procedures, Shell Scripting, XML

We'd love your feedback!