Advisory Technology Architect - Spark And Kafka Resume
Dover New, HampshirE
TECHNICAL SKILLS:
Big Data Ecosystem/Programming Languages:Hadoop 2.7, HBASE. HDFS, Flume, Scoop, Oozie, Yarn Release 2.0, Mesos, Storm, Kafka, Spark 1.6.1(SQL/Core/GraphX/MLLib/DataFrames/RDDs), Neo4J, MongoDB, Cassandra 2.0, Riak, Talend Open Studio 5.6, Pig, Hive, lambda architecture - batch versus near real-time, Aerospike RDBMS ETL, Cassandra 2.0, Riak, MongoDB, R, Java 7/8/Eclipse - Helios/Juno/Mars, Python 2.6/3.3, AeroSpike; Avro/Parquet/JSON data formats, Kerberos 5, Java cryptography (AES-256 bit), Java 8 and RestFul API; Tachyon 0.8.0, Apache Zeppelin incubator 0.5.5; Flink 0.10.0 (Hadoop 2.7/Scala 2.11 IDE), compression strategies - Gzip, LZO, Snappy, serialization - Java/Kyro, akka/Actor, Zeppelin notebooks: CrateDB/Grafana; Kavlev Data Modeling Tool(Cassandra 3.0), DataDog, Zookeeper/Kafka tuning
Cloud Technologies (AWS):EC2/EC2 Container Service(Docker), Elastic Beanstalk, Lambda(event driven architecture); Storage - S3, Elastic File System, Direct Connect, Route 53, CloudWatch(autoscaling), CloudFormation, CloudTrail, Config, OpWorks, Identity and Access Management, Inspector, EMR, Kinesis, Machine Learning algorithms, API Gateway, AppStream, enterprise AWS cloud design patterns, AWS (IoT); RedShift, Storage model analysis - hourly, reserved, spot-pricing, autoscaling
Dev Ops:Splunk, Nagios, Ganglia, JVM Tuning, MemCache, Tachymon, Project Tungsten, scripting - bash, awk, sed, shells, Linux kernel / io device modifications, Puppet/Chef/Shellshock, Ansible Tower, Linux - Red Hat RHE 6, CentOS, Ubuntu 12.3, Fedora 15/16/17; RANCID, Cacti(Graph), lldpd, IPerf, MultiHost SSHWrapper, Jenkins, Hudson CI, Bluepill, Capistrano, Bcfg2, Supervisor, Graylog, runit, Squid, snort, system, netstat, iostat, vmstat, ltrace, strace, ftrace, perf, tcpdump, sar, man, Takapi, GitHub, ssh, Cygwin, WinDiff, putty, Maven, Ant, DTraceToolkit(Netflix), Selenium, Node.js, putty socket.io/WebSockets, Azul Technologies - Zing; Stackio/Stacki(large scale Linux cluster automated deployments)
PROFESSIONAL EXPERIENCE:
Advisory Technology Architect - Spark and Kafka
Confidential, Dover, New Hampshire
Responsibilities:
- Led screening effort to pick best architects in the industry for building out the associate Big DataSolutions architect role - including Walt Disney, Confidential and Confidential ; looked for candidates
- With deep knowledge in data streaming with Kafka and Flink; created a 50 point skills evaluation and
- Big Data use case and design scenario for evaluating candidate strengths and weaknesses
- Conceived and designed a 80 page Kafka playbook for operations and tuning high performance low latency messaging systems - tuning broker, producer and consumer nodes, Java gc()
- Created a 110 page cookbook on using DataDog and creating customer performance dashboards for Zookeeper; created a 100 page cookbook for creating Kafka performance dashboards
- Conceived and created 30 page architecture design document for integrating Kafka consumer to a Hadoop batch gateway via HortonWorks release 2.4 Ambari REST services; integrated Kafka consumer with SPENEGO/Kerberos 5 authentication services; custom document for Kafka - Hadoop data transfer and integration
- Installed the Confluent 3.2.1 Kafka broker and Operation Dashboard; ran and modified the performance
- Kafka load against DataDog Zookeeper and Kafka dashboards; custom Docker dashboards for 500+
- Docker containers in Red Hat Linux 7.3 kernel; integrated Kubernetes as orchestration engine for Docker containers
- Top analysis and deisgn of the Kafka Broker, Kafka Producer, Kafka Consumer, Schema Registry(100
- Schemas), Kafka Connect for Cassandra, Hadoop HDFS; utilized the API for release 0.11 for the producers and consumers; set up the replication factor, logical compensation for automated consumer rebalancing; utilized Mirror Maker for setting up DR between two data centers; established the JVM criteria for the
- Kafka broker, producers and consumers; set up the Kafka metrics using JProbe; Confluent Control Center
- Release 3.1; devised durable commit markers for Kafka consumers - custom Java code for Kafka consumers; designed optimal consumer groups based on use case, topic type and number of Kafka brokers; conducted the Netflix Chaos site reliability engineering tests on the dev platform; drop partitions, brokers, consumers; setup up the Zookeeper cluster 3 and 5 nodes- utilized the 4 letter commands for systems admin for Zookeeper
Enterprise Data Lake/AWS Architect
Confidential, Basking Ridge, New Jersey
Responsibilities:
- Conceived and designed custom POCs using Kafka 0.10 and the Twitter Stream in standalone mode; architectedthe front-end near real-time data pub/sub non-blocking messaging system using the Kafka/Confluent.io Enterprise
- Platform; configured the 10 nodes - 3 Web servers, 4 Kafka brokers and 3 Kafka consumers(Spark Strersming(DataFrames) with 3 Zookeeper nodes; Kafka brokers able to sustain 1 million wirtes per second peak period for proprietary IoT device analytics plafotom for 4G LTE KI indicator(over 200); researched and codified the Kafka Consumer using KafkaConsumer API 0.10 and KafkaProducer API 0.10(Java); designed the Spark Streaming and KafkaProducer interfaces - for multithreaded partitions and multiple topics by smartphone manufacturer device type; competitive analysis of Storm, Spark, Flink, Samza for processing messages(once only), replay and lost message management, horizontal scalability, security, message sequencing; coordinated Kafka operationa and monitoring(via JMX) with dev ops personnel; formulated balancing leadership strategies and impact of producer and consumer message(topic) consumption to prevent overruns; aggressive monitoring of partitioning versus topic production via JMX interface(s); developed Kafka standalone
- POC’s with the Confluent Schema Registry, Rest Proxy, Kafka Connectors for Cassandra and HDFS(Hadoop 2.0);
- Custom Kafka broker design to reduce message retention from default 7 day retention to 30 minute retention - architected a light weight Kafka broker
- Created custom test, design and production Spark clusters for the VERUCA - Verizon Universal Communications
- Architecture - Spark clusters exclusively from the AWS Management Console - configuration details, network configuration and security details; architected the S3/EMRFS file systems(11 9's) for the proprietary
- Datasets for 4G LTE analytics for radio signal loss, cellular tower placement, latitude and logitude, 100+ device modem metrics from 30 million devices and 1000 central switches; Spark clusters 1.6.1 in the various environment: s; wrote custom and custom DStreams in Scala for in-flight versus at-rest data for lambda architecture; set up YARN with dynamic allocation for horizontal scaling; calculated different pricing models for reserved, hourly versus spot-pricing; configured EMR for M3/M4 AMI machines for smaller test/ development Spark cluster(8 nodes); separation of computer versus storage AWS frameworks; designed a persistent versus transient architecture - raw Linux server with Spark ML algorithm jobs, test Spark jobs via Zeppelin notebooks; mentored and guided offshore team in troubleshooting and fine tuning Spar
- SQL applications with Ganglia - server load distribution, Spark UI - cached partitions, CloudWatch console metrics, heuristic search through log files of the Spark executors on each Spark worker node and Spark driver; performance tuning of the number of cores, number executors, amount of memory and network bandwidth; code reviews for the optimal Spark application programming; analysis of DAG diagrams for Spark internal execution -
- "lazy" transformations versus actions; examination of Spark UI for job completions, job task completions, cached versus persisted DataFrames; assisted/advised the resident data scientists to configure and codify ML sets and modeling; caching strategies for multi-pass algorithms for better throughput and performance - forest clustering, finite difference calculations, normal distributions, Bayesian statistical modeling; utilized splitable compression to increase throughput from S3 to EC2
- Assisted client in technical interviews of over 20+ potential Big Data architects, technical background checks and review of CVs and resumes; set up Spark coding tests in Scala and Java; installed a Scala IDE test environment to evaluate functional coding precepts; assisted client to interview 10+ dev ops engineers and several Scala developers in knowledge of Linux operations, bash coding skills, troubleshooting techniques and scenario diagnostics
- Collaborated and advised the resident data scientists to extrapolate use cases for machine learning; designed POC
- For Spark applications written in Scala utilize the MLlib - regression, experiments in recommendation engines based on 4G KPI indicators; established performance sandboxes on signal propagation, theoretical versus actual;
- Spark applications examined hundreds of HDFS 5 Mb files across 5 million device sample; derived starndard versus normal versus Poisson distribution models cross-correlated with the cell phone tower lat/long positions across the domestic US;
- Architected full life cycle the Veruca Cassandra ring - developed and designed the entire technology stack, versaw and reviewed the APO for .75 mil for the 30 node ring for prod/qa, 15 node ring for dev; orchestrated the hardware procurement of the 1 Pb analytics data store ingested through the Kafka pub/sub cluster consisting of
- 10 nodes, peak volume of JSON data coming from 30 million smart phones(android), KPI payload of 3000 bytes/minute; replication factor (X 3); set the DataStax OpCenter for Cassandrta node analysis/troubleshooting;
- Each node consisting of 256 GB RAM, 2 Intel chips each with 8 cores, 6 spindles of 6 Tb/7200 rpm SATA drives
- (JBOD, non-RAID), 2 GB of solid state memory for RHEL version7.2, JRE/JVM rel 1.8.0 92, 10 GigE network with 42U racks, Liebert 440 UPS electrical subsystems, calculated the BTU and heat dissipation and cooling requirements with building HVAC engineers; worked with the radio/telephony 4G engineers to determine ptimum query patterns for time-series analysis for 1, 15, 30, 60 minute intervals; query patterns involved denormalization of LTE signal tables and device KPIs consisting of 200+ KPI indicators(LAT/LONG coordinates); established and formulated best practices on Cassandra design patterns - atomic distributed counter service, needle in the haystack, anti-patterns; conducted due diligence on performance tuning, read/write consistency of one/QUORUM, schema design; set up bash scripts to centralize Log4J logs by data node; vnode key distribution; synchronized all cassandra.yaml configuration files; oversaw partitioning, secondary indexes, CQL types, use of supercolumns
- Designed and architected the HA solution for Cassandra rings between Basking Ridge, NJ and Dallas, Texas for the 30 node ring at each location(T1 dual channel multiplexer link with backup); calculated the optimal snitch based on rack awareness and the NetworkTopologyStrategy, monitored performance in the secondary data center; compaction strategy for SSTables/memtables; designed logging/monitoring systems for the backup data center;
- Spearheaded the POCs for the AWS ecosystem via the AWS Management console, S3 buckets, security - multi-factor authentication, access keys, X.509 s, Eclipse ID plug-in. emphemeral/persistent storage options - Linux and Windows AMI instances, private subnets, designed and deployed Amazon CloudWatch, IAM, Elastic BeanStalk, AWS Simple notification; architected various cloud computing and service design patterns - snapshot, Vagrant, high availab ility - multi-server- floating IP; processing static data - private data delivery, direct storage hosting; patterns for uploading data - write proxy pattern, state ssharing, cache proxy pattern; cloud patterns for operation and maintenance - bootstrap, cloud dependency, stack deployment, weighted transition, hybrid pattern; analyzed t radeoffs for high availability of zones for fault-tolerance versus high availability; set up alarms for CloudWatch for recovery of a failed Linux server, and auto-scaling for guaranteed SLA’a for Linux servers for real-time streaming analytics via Kinesis; analyzed RTO/RDO availabilities for virtual servers for time-lapse of recovery scenarios; established a common network host naming convention with Route 53 with Class C address/VPC subnets; accessed from GitHub Chaos Monkey(Netflix) for arbitrary host/network high latencyperformance problem injections into a custom Dev Hadoop/NDFS cluster(10 nodes) with subsequent post enterprise engineering efforts to monitoring HA via Ganglia; collaborated with sr Web developers for custom
- Web applications - AWS Elastic Beanstalk with multi-container Docker financial applications
- Downloaded, configured Apache Zeppelin binaries/conf for Spark Web clients; integrated Zeppelin daemonwith Spark master node, tested and configured Web server with Spark cluster; tested Zeppelin with SparkSQL and Python clients(pluggable interpreters); tested screen sharing functionalities WebSockets, Zeppelin views from
Sr. Cassandra Architect DevOps/Cloud
Confidential, New York City, New York
Responsibilities:
- Created variation of the lambda architecture consisting of near real-time using Spark SQL; Spark cluster 1.4 consisting of 25 nodes running with 200Gb ram/24 Tb, about 1Pb of market data spanning 2000+ stocks with market ticks, number of shares traded, stock price, market ticks over 10 year period; Apache Open source version with Mesos job scheduler; developed, designed tested Spark SQL clients with Scala, PySpark and Java clients; selected best of breed in terms of time-to-deliver; created |Spark Contextx, DataFrames for Cassandra backend and
- DJ 30 versus SP 500; custom experiments with SP 500 indices with short term SP 500 futures; custom Spark applications designed with accumulators and broadcast variable to gain 4-5% in lowering network “chatter”;
- Spark cluster in dev environment benchmarked with the Google page-rank algorithm; set up benchmark based on the Daytona sort as reported by the University Of California Berkeley using 1 and 5 Tb; algorithmic comparisons f GraphX versus Neo4J of company ownership of Fortune 500 board of directors - business relationship connectivity analysis; tested Zeppelin(Spark UI) and Tachyon 0.8.0(off JVM memory management) options of
- Spark; configured master/standby servers; configured Tachyon in local machine, standalone, EC2 mode with AWS
- Vagrant plug-in, leveraged Tachyon I/O options for memory life cycle;utilized Spark Scala/Java API/Github; custom design and verification of Spark machine learning algorithms - feature extraction, pipelining, regression analysis, dimensionality reduction (PCA and SVD), k-means clustering
- Comprehensive design, discovery, analysis of the SP Capital IQ software, infrastructure, analytics, hardware in conjunction with the internal architecture review board - concerns of duplicate service calls, improvement and enhancement of existing SLAs to determine, document inaccurate stock quotes and improvements in real-time calculations from the legacy Soalris 9 Unix servers(200+); established comprehensive migration plan to a
- Red Hat Linux(100+) server infrastructure, incorporating complete software stack redesign; collaborated with the
- EA review board for establishing a IQSF(Intelligence Quotient Service Framework) to cover all mutual fund bondequity instruments for corporate, munis, government fixed income instruments via a SOA REST API;
- corporate wide standard of securing customer services for market quotes; comprehensive review, modification a and enhancement of over 500 SOAP service calls to REST API service calls; established and created SOA service call directory(on-line) for bid/ask/rate spreads for commodities – gold, silver, platinum, palladium futures, Forex 30/60/90/120/360 for over 100 currencies; assisted peer architect for identifying use cases for Riak and MongoDB annual reports, filings 10K with SEC for the NASDAQ and NYSE for 5 year span, 1.2 M pages in Adobe text, searchable by financial keyword – asset, liability, receivable, payable, shares of stock
- Successful integration of Cassandra 2.0 distributed logger; very high volume – supports the S&P 500/Dow Jones industrial indices;over 20+ nodes integral market data infrastructure support the SP Capital IQ real time desktop global delivery system; Cassandra Ring has DataStax Enterprise Edition, replete with OP Center; installed
- Managed, configured, tuned and continuous deployment of 80 Hadoop nodes in a Red Hat Enterprise edition 5; configured via the AWS console for 2 medium scale AMI instances for the Name Nodess, 78 large scale Data Nodes with 8 Intel i5 cores,3.5 Tb of disk and 350 Mb for JVM per Data Node; automated deployment and Linux system configuration via Chef; utilized 25 different dev op tools to log, debug, discern diagnose performance problems at the database level, Linux daemon level, networking level; set up real-time alerts with custom scripting via awk/fgrep/grep for kernel thread utilization; JVM tuning and garbage collection of short versus long lived Java objects on different generation heap spaces with due diligence on “stop the world gc() algorithms, “mark and sweep”; Chef automated deployment on qa Hadoop cluster of 80 nodes (mirror of prod Hadoop cluster); deployment and configuration of 20 Hadoop nodes on AWS AMI Linux instances
Corporate Security Expert
Confidential, Harrisonburg, Virginia
Responsibilities:
- Comprehensive review and analysis with a complete top down assessment of corporate records retention, storage, destruction policies; complete review of all infrastructure artifacts – databases, middleware, firewalls, DMZ,\ network routers, subnets, honeypots, SSO/LDAP configurations, hardening and rotation policies of corporate and external users of rosettastone.com, Web/Apache server/Ubuntu 11/12 kernel hardening/patch reinforcements; top down review, design and rollout of 3 million customer Visa and Confidential numbers state-of-art encryption strategies – two keyTriple DES, Skipjack, NIST/NSA advanced encryption standards and recommendations; review of all corporate email systems for virus and SPAM control, revised strategies and techniques for external countries for currency exchange, foreign payments and auditing, field activity reporting; instituted quarterly ethical hacking procedures, reporting, analysis and follow up IT engineering endeavors; including establishing a corporate security lab to test the latest in pen tests for Windows, Linux
- MacOS and Android/Apple smartphones and tablets; instituted a corporate wide systems responsibility and charter for hardening 4000 company laptops for common encryption/decryption procedures to prevent internal software program theft; instituted and rollout of Kerberos 5 for internal security/ticketing for all J2ee applications running JBOSS 6/6/1 cluster servers for QA and production environments
Enterprise Design Architect
Confidential, Minneapolis, MN
Responsibilities:
- Launched and promulgated custom business rules engine framework, consolidated and interviewed key
- SME’s on pharmaceutical rules and medical conditions based on the National Drug database and
- Comprehensive review of all retail insurance process artifacts, rules engines, message buses, business transformation models, security enforcement of HIPPA /HL7 relating to scrubbing Confidential t data, review of over 5000 + insurance policy due diligence of health and sickness criteria; developed the Aetna Comprehensive Insurance Screening Framework(ACMSF)based upon the precursor of the Affordable Health Care Act; integration of the Kerberos 5 authentication and adjudication policy audit server tracking 3 mil+ inquiries into PPO/HMO/Medicare customers; ACISF built according to the TOGAF 9 methodology; bi-weekly meetings with key executives and stakeholders from the Aetna Enterprise Architecture Review Board for reporting and software and infrastructure component resilience, security, fault-tolerance, performance metrics and SLA’s (4 month effort with business constituencies) resulting in 250 pages of schematics with a 10 EAF steering committee; successful integration into Tibco and Websphere SOA Orchestration server; extensive utilization of best practices of various enterprise integration design patterns for message proxy, modified “spoke and wheel” topology for QA and production messaging frameworks across corporate messaging bus; integrated REST service APIs(over 400+) service calls for insurance policy look ups, claim processing, special APIs created for high speed lookups for insurance actuary tables(Gigaspaces XA) in-memory cache
Credit Default Swaps Trading Architect
Confidential, Pennington, NJ
Responsibilities:
- Application(s): FIX 4.5, credit default swaps, fixed income trading, Dodd-Frank compliance; Monte Carlo risk analysis/payoff matrix scenarios; HFT custom algorithms analysis and design; custom Java/C++ software for multiple precision routines(up to several hundred places) interest calculations, factorial/Fibonacci series
- Technology stacks: Spring MVC/Acegi, Weblogic 10, custom Java/C++ v 1/Boost/STL software; Red Hat Linux/Solaris 10; Hudson CI/Jenkins/Maven; DB2 UDB 8.0; custom design pattern(s); PVCS; JProbe; Python/Jython scripting; Gemfire data caching; SAML 2.0/SSO/Java private/public key cryptography/X.509 digital /passkey/passphrase generation/management/”honeypot” DMZs
